CHEM455/CHEM555: Chemometrics

Table of Contents

Class information

First day lecture and some motivational words

What is Chemometrics?

My Definition of Chemometrics

The application of numerical methods to extract chemical information from chemical data.

What are numerical methods?

Math

What is chemical Data?

numbers that you get from measurements

What is chemical information

  • results
  • constants
    • equilibrium constants
    • rate constants
    • bond lengths
  • knowledge
  • models
  • conclusions??
  • etc.

Examples of Chemometrics: Concentration Determination.

All analytical chemistry deals with mixture analysis. From this perspective there are two issues that are coupled.

  • What are the compounds in the sample?
  • How much of each are there?
  • This first part deals with "How much is there?"

Determining concentration of an unknown sample.

Let us pretend you are a chemist and need to know the concentration of phosphate in an aqueous solution. What do you do?

  • First is the calibration standards (external calibration)
    • prepare a series of phosphate solutions of known concentrations
    • measure the UV absorbance spectrum of each solution.
      • WAIT A MINUTE!!!!
      • you don't measure absorbance, what do you measure?
      • if you are using a photocurrent detector for your measurement,
        • then you measure current as a function of wavelength (sometimes called a single beam spectrum) of each solution including from a blank sample
        • typically the computer software of you spectrometer converts single beam spectra of each solution and the blank into an absorbance spectrum using the following equation
    \begin{equation} A = -log_{10}\frac{I_s}{I_b} \end{equation}
  • Absorbance Spectra of Phosphate

    beerslawdata.png

    Figure 1: Absorbance spectra of sodium phosphate solutions

  • Create table of concentrations of phosphate and band absorbance
    Table 1: Phosphate data
    Concentration of Phospate /mM Absorbance at 210 nm
    0.100000 0.097067
    0.300000 0.296739
    0.500000 0.468451
    0.700000 0.669580
    0.900000 0.861085

    Note: 210 nm is the top of the spectral band for phosphate. This is the analytical wavelength.

  • Make a Beer's Law Graph

    beerslawgraph.png

    Figure 2: Beer's Law graph of phosphate absorbance at 210 nm vs. known concentration of phosphate

  • Fit the data with a linear function

    Because of the linear relationship between concentration of photon absorbing species and the number of photons absorbed. (This relationship is called Beer's Law.)

    \begin{equation} A = \epsilon l C \end{equation}

    we can fit the data with a linear function.

    \begin{equation} A = m C + b \end{equation}

    This equation (specifically the slope and intercept) is a model for the behavior of this chemical system (chemical information).

    • m is the slope (sensitivity of the photon absorbing analytes to absorb light)
    • b is the y-intercept (empirical parameter that encapsulates the optical differences between the blank and the calibration standards, ideally this is zero.)
    • these (m and b) are also called the fitting parameters

"Measure" the aborbance spectrum of the solution of unknown concentration

  • and get its absorbance at 210 nm. Remember this wavelength (210 nm) is the analytical wavelength, because we are using the absorbance at this wavelength to determine concentration.
    • in this case it is also the λmax, which is the wavelength of maximum absorbance
  • solve the linear function for C
\begin{equation} C_\text{unk} = \frac{A_\text{unk}-b}{m} \end{equation}
  • and plug the unknown solution's absorbance at 210 nm into the rearranged linear function
\begin{equation} C_\text{unk} = \frac{A_{\text{unk}}-b}{m} \end{equation}

Cunk is the concentration of phosphate in the solution, and is also chemical information.

What assumptions were made during this example?

  • unknown solution has the same chemical environment as the calibration standards
  • interactions between absorbing species do not change with concentration
    • usually means: relatively low concentrations (therefore we can ignore Activity)
  • no interaction between absorbing species and matrix
    • usually means: don't have to use method of standard addition
  • the analyte (phosphate in the example above) is the only absorbing species at the analytical wavelength
    • Very simple system,
    • requires lots of unknown sample preparation to get the solution to only contain phosphate
    • lots of human labor for highly variable situations (blood, river water, detergent factory run-off, drainage near a mining operation, etc.)
    • and therefore of only limited use
    • or involves chromatography to cleanup the sample and takes a "long" time per sample

In the example above, the sample and calibration standards contain phosphate PO43+ + Na+ in water. It is a mixture, but only the phosphate anion absorbs UV light. So from a spectroscopic/chemometric perspective this chemical system is a single component system.

How do you deal with mixtures and Beer's law?

Another word for mixture analysis is multicomponant analysis. There are 2 situations for calibration curve creation and unknown concentration determination.

1: You know every component's pure spectrum and concentrations

A B
spectra don't overlap spectra do overlap
can use 2 curves ???
  • Example: 1.A

    mixspecsep.png

  • Example: 1.B

    mixspecol.png

    mixspecol_series.png

    Figure 4: Series of mixture spectra

    mixspecol_bl.png

    Figure 5: Beer's law graph of the two components

    In this situation, Components' A and B spectral features overlap. How do you deal with this? We will learn some strategies for this situation this semester.

2: You do not know every component's pure spectrum nor do you know all of the concentrations of all of the species that contribute to the spectra

In other words, you know the concentration of your analyte in the calibration standards but not the concentrations of other spectral contributors in you unknown samples, and maybe even your calibration standards.

This is a bad situation that is all too common in the real world. You might be asking yourself, why does this situation occur? Here are some reasons:

  • the sample that you need to analyze is precious (think Gollum from the Hobbit and Lord of the Rings) and can not be replaced or damaged in the measurement. Examples
  • the sample as a mixture has important value that is greater than the purified parts (think BBQ sauce, or steak)

Some examples that are real but don't fit into the categories above

  • NIR spectroscopic determination of glucose concentration in blood (everything in the blood contributes to the NIR spectra of blood)
  • but for a diabetic knowing the glucose concentration in their blood is critically important for maintaining their health
  • UV spectroscopic determination of nitrate in natural waters (nitrate, phospate, and chloride all "look" similar to each other in the UV spectra)

simpmultispec.png

Figure 6: Toluene and p-xylene UV absorbance spectrum in a series of 2 component mixtures

You will learn in this course, how to predict the concentration of multiple components in mixtures of both situations 1 and 2.

Examples of Chemometrics: What is it?

All analytical chemistry deals with mixture analysis. From this perspective there are two issues that are coupled. What are the compounds in the sample, and how much of each are there. We saw some examples of the second part of this sentence earlier. These types of tasks are sometimes called quantitative analysis. This first part of the sentence deals with "What is this unknown material, and how does it relate to other materials?" In other courses this task is sometimes called qualitative analysis.

Identification of unknown material: IR spectral classification

In organic chemistry you probably did some IR spectral analysis in which you identified parts of a compound by its IR spectral features. For example, you might have been given the following IR spectrum and asked to determine if it contains a carbonyl functional group or not, and if so what kind of carbonyl (i.e. amide, ketone, carboxylic acid, aldehyde, etc)

exidspec.png

Figure 7: What kind of molecule is it?

What is it?

Lots of analytical tools will give us data that can be used to determine chemical identity.

  • lots of examples at WCU (and elsewhere) in undergrad classes
    • MS
    • NMR
    • Raman
    • IR
    • GC
    • TLC

Example questions from Forensic Science

  • What polymer was used in this textile?
  • Is this powder cocaine?
  • Is this powder an explosive residue?
  • Is this powder anthrax spores?

These things can be done by visual inspection of the data. An example is spectral interpretation, and manual comparison of bands/peaks in a spectrum to a set of reference tables. These both take lots of time.

  • Washington, D.C. anthrax spores were sent in the mail.
  • Ideally: need to scan every letter that passes through the mail system (millions per day)
  • Whenever we need to repeat a task many, many times, often it is better to let a computer do it.
  • The trick is figuring out how to train the computer to do what a human can do.
  • This is what chemometrics is all about.
  • Easy for a few not so easy for thousands

    figureprint.png

    Figure 8: Finger print, DNA, RDX

Chemometrics is

  • getting the computer to process your data for you, rather than processing it by hand.

Therefore we will need a computer and software to do this.

Industry standard options

  • Canned software (operating system dependent)
    • ThermoFisher Grams
    • CAMO's unscrambler
    • Others
  • Matlab computer language (operating system independent)
    • very nice
    • a lot of DIY (therefore good for learning)
    • toolboxes just for chemometrics
    • additional price for toolboxes
    • on school computers (not toolboxes)

Why not MS Excel? So you noticed that MS Excel was not listed above as a tool for doing chemometrics?

  • you are a scientist
    • MS Excel is a tool for buisness people who are afraid of doing math
  • I am a MS bigot
  • spreadsheets are not good for massive amounts of data
  • So this is a choice. You can make your own, but I am going to use another choice for this course.

Octave is a free choice

  • similar to MATLAB (mostly compatable with Matlab)
  • free
  • open source
    • if your really like it you can contribute to it or help maintain its awesomeness

R, S, SAS, SPSS

  • relatively easy to use
  • these are called area specific languages (they do stats well, but only stats)
  • little difficult to understand what is happening underneath the hood
  • I prefer something else, and I am the teacher

Python is a modern choice

  • One limitation of all of these other choices is that they are all area specific tools
  • this can be a limitation, if you ever want to do something else besides process data
  • Python is a general purpose computer language
    • it can be used for most computer programming tasks
    • it is becoming, if not already is, the language for Data Science
  • Chemometrics is a subset of Data Science (and pre-dates it by more than 40 years)
  • this is what we will use for this class

Getting the software and assignments

Anaconda distribution

We will use the Anaconda distribution for getting python in this course. Some benefits are

  • all self-contained
  • easy to install all of the parts
  • does not interfere with the rest of your computer

You have 2 choices:

  • use the http://virtual.wcu.edu (requires internet/network access to use, but very little computer resources)
  • install anaconda
    • goto https://www.anaconda.com/download/
    • download the Anaconda installer for your computer (Linux, MacOS, Windoze) Python version should be > 3.6 (Python 3.9 as of 2022)
    • install it on your computer
    • run it to make sure that it is working

Here is a pdf with screen captures of me installing the Anaconda distribution on a Windoze computer: ./figures/anaconda_installation_running_guide.pdf

ass01

Prove to me that you able to run python. Go to your menu system and start Jupyter notebook. This will open a tab in your default browser. In Windoze, it should look something like:

./figures/anaconda_installation_running_guide.pdf

  • For now, select Python 3
  • This should open a new tab in you browser, running jupyter/python
  • Do the following in Jupyter Notebook
Assignment 1:  1 point
In the Cell beside In [ ] 

Type the following things
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

and press Enter on your keyboard while holding down the SHIFT key

print('Your Name')

and press Enter on your keyboard while holding down the SHIFT key

Screen capture the Juypter tab.
Upload this image file in the Learning Management System DropHole under ass01 or "Proof of Installation"

  • You can delete this Jupyter notebook file after you email the image.

Dropbox hole

In this course, we will use dropbox to share files. So your second assignment is to get a Dropbox account (if you already have one you can use that one). With this account is an associated email address. It may be a catamount email address or it may be something else. Send me an email from this dropbox associated account. Put in the Subject line:

Assignment 2: Dropbox

I will then share a dropbox folder with you. Once my account and your account are linked in this directory, you should see in that directory/folder the course files that we are sharing. Note: I will not be able to see any other dropbox stuff of yours but only this one directory. We will refer to directory as the course home directory. Make sure that when you run Jupyter that you can access this course home directory. (One way to do this is to run Jupyter notebook/lab from this directory.) Another way is to put the Dropbox directory within your course directory on your computer. See the following structure.

locdropboxexample.png

Figure 9: Example of where to put your Dropbox on a Windoze computer

chemoclass
├── Dropbox
    ├── learn
    ├──resources
    ├──ass01

In this course home directory will be several subdirectories:

  • resources - this is the textbook, and papers to read
  • learn - this is where you can practice your stuff
  • assignments - there will be lots of these that correspond to the assignments. Each assignment will have a unique directory. They will begin with:
    • ass## - assignment directory the ## is 01, 02, etc that corresponds to the assignment number

An example homework directory may already be in the dropbox course home directory, called ps1. You can poke around there if you want. This directory will never be accessed by your instructor.

Here is what the course home directory will look like.

├── learn   <-- practice stuff here
├── ps1    <-- example assignment
└── resources     <-- reading material (the text book in pdf form)

The learn directory

The learn directory will have the following structure at the start. This is were you can do most of your learning/practicing for this course.

learn
├── data
├── figures
├── lib
└── notebook_template.ipynb

The whole enchilada

Here is what your dropbox course home should look like the at the beginning of this course.

├── learn
│   ├── data                       <-- this is where learning/practicing data that you will process is located
│   ├── figures                    <-- this is where you save your figures generated in python 
│   ├── lib                        <-- in this directory are library files that we will use later
│   │   ├── analtool3.py
│   │   └── chemo.py
│   └── notebook_template.ipynb
├── ps1                           <-- practice assignment 
│   ├── jupyter.png
│   ├── problem1.ipynb
│   └── problem2.ipynb
└── resources
    └── textbook.pdf

Assignments

Whenever you work on an assignment, make sure that you don't change the name of the files, and make sure that the files are in the same directory from which they came.

Once the assignment has beed graded, you will receive feedback in the same directory as the assignment and/or via email.

ass02

Assignment 02 will be proof that you have installed dropbox on your computer and can receive an assignment and submit it back to me. In the ass02 directory open the jupyter notebook and type your name in the cell below the instructions.

Introduction to Python in Jupyter

Here is a tutorial for python. Much of the ideas in this learning module can be found at this tutorial.

https://docs.python.org/3/tutorial/

Introduction to Jupyter

Whenever you run Jupyter on your computer, you will see in your browser the following:

jupyter_home.png

From this screen you can either start a new python notebook or open an existing one. In your home/learn directory you will find at least one notebook template. Copy that template to you working Jupyter directory. From the home screen you can check the notebook template checkbox and an option at the top to duplicate the notebook will appear. Do this. The reasoning is that you preserve the template for future use and only work on copies of it. You can rename the copy to anything that you like.

jupyter_whole_screen.png

Figure 10: Jupyter notebook page running

To run python code, you type the code in to an active cell in Jupyter (as indicated by the blue bar on the left side of the active cell). Once you have finished typing the code to execute (or run) the code you simultaneously depress the Shift-Enter keys. (hold SHIFT and press ENTER)

Use the template

A lot of our code in this course will rely upon pieces of code that other people have written, these pieces of code are stored in modules that come with the Anaconda python distribution that you have hopefully already installed. A template has been provided in your learn directory. You can start with this template and make a copy of it directly from Jupyter. See if you can figure out how. I would suggest that you keep a pristine copy of this template. That way if you break what you are working on, you can go back.

learn
├── data
├── figures
├── lib
└── notebook_template.ipynb <------HERE


A computer is a calculator

For the purposes of this course and in my opinion, a computer is nothing more than a calculator (although a big, powerful, awesome one). (parts 1-5)

What do you do with a calculator?

  • read/write data
  • process data
  • plot data
  • evaluate data
  • extract information from data (math)

What are the parts of the calculator

  • user interface for input (for us: jupyter)
  • user interface for output (for us: jupyter)
  • a computational engine for doing the math (for us: python)
  • storage space for saving things for later usage (for us: jupyter for temporary and your hardrive for longer term)

comments in python

  • Use the # symbol to add comments (extra info that might be useful latter but is not used in the calculations) to your math or code
  • definition: code are the words and other groups of letters and numbers that you type to get the computer to do work
  • Review: make a comment in a cell and execute the cell. See that the Python compute engine (kernel) ignores it.

Some Rules of python

  • the code in python is case sensitive, so a variable named tup is different than Tup, and TUp and TUP, etc.
  • some symbols have specific predefined meaning in python:
    • all variables must begin with either a letter or an underscore
    • after the first character the remainder of your variables may consist of any number of letters, numbers or underscores
  • Don't use the following words in variable names that you create
and assert break class continue
def del elif else except
exec finally for from global
if import in is lambda
not or pass print raise
return try while    
False class finally is return
None continue for lambda try
True def from nonlocal while
and del global not with
as elif if or yield
assert else import pass  
break except in raise  

or these

numpy Float Int Numeric Oxphys
array close float int input
open range type write zeros

or these

acos asin atan cos e
exp fabs floor log log10
pi sin sqrt tan  
  • white space matters: don't use spaces in filenames or functions or variables
  • don't use hyphens - or parentheses or brackets ( or ), [, ], {, or ] in variable names
  • don't use any special character other than underscores invariables names

All of these special characters have specific meaning in python and will be mis-interpreted

  • Review:
    • type some words from the tables above in the cells of jupyter to see that the change color (this is called syntax highlighting), and it means that Jupyter recognized those words as predefined Python words
    • type some variables in a different cell disobeying the rules (just do one at a time) and execute the cell (Shift-ENTER or Shift-Return). Look at the error messages that are created.
      • Use this as a template for defining a variable (x is the variable).
    x = 1.0
    
    

displaying stuff

  • use print( whatever you want to print, separated by commas) to print stuff to the screen
  • Review:
    • use the print function to communicate to the user of the python code. Make your code print the value of a variable.

arithmetic

Computers can do arithmetic. Here are the commands for various arithmetic operations.

Type the following things in a jupyter cell and press shift-enter after each one:

print(4 + 4)  # this command prints the results of 4+4 to the screen
4+4           #this one does not (unless it is the last command in the cell)
print(4 - 4)  #What does this do?
print(4 * 4 ) #and this
print(4 / 4)
print(4**2)


8
0
16
1.0
16

in these examples the +, -, , /, and * are called operators.

Note: In Python ** is the operator symbol for power.

  • Review:
    • Use Python to solve a chemistry problem such as determining the number of moles in a 0.08841 g of NaCl (Just type in the numbers and operators.) and print the value the screen.

algebra

Lets say we have this simple algebra equation

\begin{equation} 4=x-1 \end{equation}

and your task is to solve for x. What do you do? You subtract -1 from both sides to give you

\begin{equation} x = 4+1 \end{equation}

in this equation x is called a variable. In a cell, type:

x =5      #this command assigns the value 5 to the variable x
          # but will not print the contents of x
print(x)  #this command prints the contents of x
5

Now type in a cell

x =  4+1   # this command assigns the results of the operation 4+1 to x 
           # but will not print the contents of x
print(x)   # ths command prints the value contained in x
5

Notice how you get the same results. This means that the computational engine (python interpreter) is computing 4+1 and storing it into the variable x.

Note for later: 5 is called a scaler number.

Once you have this variable filled you can perform operations on it. Type

x = 5
print(x*5)  ##this command prints x times 5 to the screen


25

or we store this result in a new variable

x = 5
u = x*5

print(x,u)  #this command prints x and u to the screen
5 25
  • Review:
    • Use Python to solve a chemistry problem such as determining the number of moles in a 0.08841 g of NaCl. Use variables to store the mass and Avagadro's Number, print the value the screen and store the answer into a new variable.

Fancy printing part 1

x = 5
u = x*5

print(x,u)    #this command prints x and u to the screen
print('x*5 = ',x)  #this command prints x to the screen with some text showing how it was calculated
5 25
x*5 =  5

Fancy printing part 2

The previous output may not necessarily be what you want. Take this example:

# 
pi = 3.141592653589793238462643383279502884197169399375105820


print('pi = ',pi)  # this prints an undefined by you amount of digits
print('pi = {:.2f}'.format(pi))  # in this case you choose the significant digits to print

pi =  3.141592653589793
pi = 3.14

This example uses the format function. It does 2 things.

  1. it converts the number stored in the pi variable to a string (letters and numbers and symbols) for display/printing purposes
  2. it allows you to control what that string looks like

Why do this? Humans and computers need different formatting to read information. This format command converts computer looking information to human useful formatting.

In the second print command the f tells the computer that the number (pi) is a f*loating point number (has a decimal point). The *.2 tells the computer that you want 2 digits to the right of the decimal.

Here is a link for how to control this formatting. You can explore this at your leisure.

https://docs.python.org/3.4/library/string.html#formatspec

  • Review:
    • Store you name as a string variable
    • Figure out how old you are in Days (http://jalu.ch/coding/days/en)
    • Store this number in a variable
    • Calculate the percentage of your living days in a century of days ( you can use 36525 as an approximation) and store it as a variable.
    • print a sentence (string using single quotations) that says something like: My name is X, and I have lived Y percent of a century.
      • Use the format method (command) to substitute your name and century percentage in the printed sentence

geometry

Given the figure below what is the distance r?

geometryfig.png

Figure 11: Distance in 2D space

You may or may not remember the Pythagorean theorem

\begin{equation} r^2 = a^2 + b^2 \end{equation}

where

\begin{equation} a = 3 -1 \end{equation}

and

\begin{equation} b = 2-4 \end{equation}

To solve for r, we need to take the square root of each side of the equation above.

\begin{equation} r = \sqrt{a^2 + b^2} \end{equation}

Let us calculate r using python. Type

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np     #this command loads all of the functions in numpy and labels them np
a = 3-1
b = 2-4
r = np.sqrt( a**2 + b**2 )  ## note the np.sqrt, note the ** for power
print('np.sqrt( a**2 + b**2 ) = ',r)
print('np.sqrt( a**2 + b**2)= {:2.2f}'.format(r))
np.sqrt( a**2 + b**2 ) =  2.8284271247461903
np.sqrt( a**2 + b**2)= 2.83

What is the np.sqrt word in python? It is a built-in function; meaning that it comes with python as part of the numpy module that we loaded as np. Functions have the format

output = function( input ) Note the type of parentheses

where output (called returns in the python lingo) is the result of the function doing its thing, and input is the directions, data, and numbers that the function uses. This function np.sqrt computes the square root of a number.

Probably every mathematical function that you might need is available in the numpy module.

  • General ideas about functions
    • they can have more than one input (called parameters in the python lingo)
    • the inputs can be any kind of data structure, i.e. numbers, letters, matrices, vectors, etc.
    • to assign output to a variable simply used an equal sign =
    • function usage: outputvariable = function(input)
    • help about a function can be found by typing a ? in front of or after the function

type ?np.sqrt in a cell and execute the cell.

Review: Look at the numpy mathematical functions reference page

https://numpy.org/doc/stable/reference/routines.math.html

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions

import numpy as np     #this command loads all of the functions in numpy and labels them np

pi = np.pi                   ## notice that we also get pi from the numpy (np) module
xrounded  = np.round(pi, 5)  ### round pi to 5 decimal places 
print(xrounded)

3.14159

Functions and a few commands

Showing the cell line numbers

If you want to show the cell linenumbers, click to the left of the [ ]: of any cell. You should see a thick vertical blue bar. Then depress Shift-L. This key-stroke sequence will toggle the cell line numbers on and off.

Some more python/jupyter info

First some more ideas about how jupyter/python

  • Each time you start jupyter is called a session. By default all of the memory is wiped for each session.
  • Each time you execute a command block or cell (type the command and depress the shift-enter keys) the results are stored in the memory for that session.

type each in a cell and execute:

x = 1

who     ##this command prints the variables currently in memory

y = 100


whos   # this is a "magic" command and prints variables and their types

Building your own functions (ideal setup)

The numpy module for python contains probably every math function that you might want. Here is the website. https://numpy.org/, so you can check it out.

What if numpy or Python does not have a function that you want? What do you do?

Build you own!

For example, if we wanted to find the angle Q in the geometry problem from 11. You would need to use the following trigonometric relationship

\begin{equation} Q = arctan\left( \frac{a}{b} \right) \end{equation}

How do you go about creating a function to find Q?

First solve the problem (for now, not a function)

Before we create a function lets programmatically solve the problem, then wrap the solution into a function.

Let us solve the angle problem that we saw before and find Q. Type the following in a Jupyter cell and execute

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np


a = 3-1
b = 2-4
Q = np.arctan(a/b)
print('the angle Q is equal to {:2.3f}'.format(Q))
the angle Q is equal to -0.785

OK, so this gives you a benchmark in which to compare, i.e. the correct value for Q. Our function should give us the same answer. Note that the angle is in radians not degrees. This is the default.

OK now a function

Defining a new function in python involves wrapping the code with

def function_name( inputvariable1, inputvariable2):          ## end the first line of a code block with a colon
    '''
    this is a comment section or help section that prints when you type ? before a function in jupyter
    notice the indentation.  this is important in python
    '''
    import numpy as np

    ### this is where you type the code for your calculations

    return outputvariable  ##return stuff back to the caller

Note: In python, blocks of code are delineated by indentation

Now wrap the findq code above into a function.

def findq():  ## end the first line of a code block with a colon
    '''
    for now let's just have the function make the calculation and print it out.  
    Later we will make it more flexable
    '''
    import numpy as np                 #this command loads all of the functions in numpy and labels them np


    a = 3-1                  #Coordinates hard coded into the function 
    b = 2-4                  #Coordinates hard coded into the function
    Q = np.arctan(a/b)
    print('the angle Q is equal to {:.3f}'.format(Q))
    return
#notice the indentation is back to the first column.  So the above code block is ended
findq()     # you need the () they are part of the function.

the angle Q is equal to -0.785

Note: when you define a function you have not run or executed or called the function. To do that you have to use the function's name or call the function. (See the last line in the code block above.)

What does this function do? It calculates the angle Q from a set of fixed coordinates that are coded into the function and prints the value of Q to the screen. This approach has several limitations:

  • you need to redefine the function for a different a, b, and Q. (type in the triangle's coordinates for a and b and execute it (Shift-ENTER)
  • To use the value of Q in this form you would need to copy the value by hand and type it again (or copy->paste)
  • The value has been truncated. That could give you rounding errors further down the line of calculations. (Remember you don't round until the end of you calculations.)

Now a more useful function

Great. Now let us make this function a little more readable and useful. In the function findq

  • add some comments to the header portion of the function
  • some commented definitions to explain what this function does.
  • return the value of Q to the caller
def findq():  ## end the first line of a code block with a colon
    '''
        this function calculates the angle between to line segments.
        Usage:  angle = findq() prints the answer to the screen and returns the angle 
        The angle has units of radians
        '''
        import numpy as np                 #this command loads all of the functions in numpy and labels them np


        a = 3-1
        b = 2-4
        Q = np.arctan(a/b)
        print('the angle Q is equal to {:2.3f}'.format(Q))
        return  Q    #return angle Q 

angle = findq()                               ##store q in the variable angle
print('The square of the angle is ',angle**2, 'and the angle is',angle) #now we can use this angle for something else 

What does this function do?

  • It calculates the angle Q from a set of fixed coordinates that are coded into the function
  • prints the value of Q to the screen.
  • it returns the value of Q with all of its digits (that can be utilized in future calculations)

What are its limitations:

  • it only calculates the angle of a very specific triangle

and now an even more useful function

This function findq is not that useful; it only calculates the angle of a very specific triangle, prints the value to the screen, and outputs the angle.

You could do the same thing a lot easier without a calculation:

def sameasfindq():
    '''
    same behavior as findq
    '''
    print('the angle Q is equal to -0.785')
    return -0.785398163397
angle = sameasfindq()
print('The square of the angle is ',angle**2, 'and the angle is',angle)
the angle Q is equal to -0.785
The square of the angle is  0.6168502750673807 and the angle is -0.785398163397

It is not quite the behavior that we want. A better function would be one that could be used for multiple problems i.e. a more flexible function.

Let's alter the function to make it one that can calculate the angle given by a user's set of line segments. NOTE: Who is this user? It could be you. Or a future you. Or a collaborator to whom you share the function?

Users of the function tell function what to use for input by specifying the variable inside the () in the definition.

def findq(a,b):   ## end the first line of a code block with a colon
        '''
        this function calculates the angle between to sides of a triangle.
        Usage:  Q = findq(a,b) prints the answer to the screen and returns the angle 
        The angle given has units of radians
        a is length of the opposite  side
        b  is the length of the adjacent side
        '''
        import numpy as np                 #this command loads all of the functions in numpy and labels them np


        #a = 3-1    #commmented out so you can see the change
        #b = 2-4    #commendted out so you can see the change
        Q = np.arctan(a/b)
        print('the angle Q is equal to {:2.3f}'.format(Q))
        return  Q    #return angle Q 
angle = findq(3-1,2-4)                               ##store q in the variable angle
a = 3-1
b = 2-4
angle=findq(a,b)
print('The square of the angle is ',angle**2, 'and the angle is',angle) #now we can use this angle for something else 

the angle Q is equal to -0.785
the angle Q is equal to -0.785
The square of the angle is  0.6168502750680849 and the angle is -0.7853981633974483

This is a more flexible function.

Review:

write a function that will use the general solution to the quadratic equation

\begin{equation} 0 = ax^2+bx+c \end{equation}

Your function should do the following:

  • take 3 numbers (a, b, c from the equation above) as inputs
  • solve for the roots of the equation
  • print the roots to the screen
  • return the roots

Do this now.

Assignment 3: ass03 toy code

Do the ass03 assignment in your dropbox home.

  • There is 1 notebook.
  • There is 1 part in the notebook.
  • Due date is written in the txt file in the ass03 directory in your dropbox home

The purpose of this assignment is to practice a little more using Jupyter and to learn (both for your instructor and for you) about how to use notebooks with assignments embedded in them, and to write a function.

Some more Jupyter knowledge (Markdown)

practice making headers, lists, and math equations in juytper using the markdown cell type

  • Make the is list in Jupyter
  • typeset the following equation in Jupyter
\begin{equation} x = \frac{-b \pm \sqrt{ b^2 - 4 a c }} { 2 a } \end{equation}

User input

Sometimes when you need input from the user the easiest is to use input.

Here is an example.

ui = input('type something ')
print(ui) 
ui+1
ui = float(ui)

print(ui+2)

Control Structures: if

Now to Controlling things

Computers are not useful unless they make our lives better or easier. Consider the following situations:

  • adding up all of numbers from 1:100
  • adding up all of the even numbers from 1:100
  • determining the square root of a list of numbers
  • reading a file line by line until you reach the end of the file

Here is the first one in code form:

x = 1+2+3+4+5+6+7+8+9+10+11+12+13+14+15+16+17+18+19+20+21+22+23+24+25+26+27+28+29+30+31+32+33+34+35+36+37+38+39+40+41+42+43+44+45+46+47+48+49+50+51+52+53+54+55+56+57+58+59+60+61+62+63+64+65+66+67+68+69+70+71+72+73+74+75+76+77+78+79+80+81+82+83+84+85+86+87+88+89+90+91+92+93+94+95+96+97+98+99+100
print(x)
5050

Typing all of this is annoying, and if we have to type everything then the computer has not made our lives easier, nor have we gained any time if we want to repeat this computation, or alter it (only the odd numbers), or expand it 1 to 1000.

So most computer languages contain structures that are time saving devices or allow for the program to alter its computational trajectory if a situation changes.

In python there are 3 that are important for this course.

  • if then else
  • for loops
  • while loops

To help you understand how theses structures work, we will use children's games.

If then else

This control structure can be understood under the context of the children's game Simon Says

Rules of the game:

if (the leader says "Simon say's" before they tell you to do something):   [condition]
  then you do it                                                           [conditional action] 
else: 
  don't do it                                                              [default action]
end of turn

Translation of this previous block into code,
the text in the ( ) after the if statement is a testable condition (usually a mathematical comparison).
If this condition is True (sometimes designated as 1) then the [conditional action] code block is performed.
The else tells the computer that the default code block is coming up next If the condition is False (sometimes designated as 0) then the default code block is performed. The end lets the computer know that the control structure is over

Here is a toy example of using the if statement to see if two numbers are equal

x = 4
if x == 3:                   #in python the conditional if line ends in a colon (:)  
    print('x was 3')
    x = x**2
    print('now it is nine')
else:                       # there is also a colon after else: 
    print('x was not 3')   

                                  # in python a code block (the whole if block) is "ended" by returning back to the 
                            # same column number as the i in if
print('Thanks for playing!')
x was not 3
Thanks for playing!

NOTE:

  • the = symbol is for setting something equal
  • the == symbol is for comparison

There are many kinds of testable conditions that can go after the if statement, but some common ones are

  • == (is equal to) Note that there are 2 equal signs next to each other. This is for comparison.
  • < (is less than)
  • <= (is less than or equal to)
  • >= ( is greater than or equal to)
  • > (is greater than)
  • != not equal to

Here is another kind of if testing:

refletters = 'abcdefg'      #this is a string 
questionletter = 'g'        #here is another one

if questionletter in refletters:                            #remember to end with a :
    print(questionletter, 'was in the set',refletters)
else:                                                      #a new code block begins so remember to end with a :
    print(questionletter, 'was NOT in the set',refletters)

print('Thanks for playing')

g was in the set abcdefg
Thanks for playing

NOTE:

  • the in conditional test means just what it sounds like "in"
  • Also try if questionletter not in refletters and see what that does.
  • Also try if questionletter == refletters and see what that does.
  • <
  • >
  • ~=

These logical notations work as well but for strings may not be as easy to understand. The in and not use in these conditional tests are examples of how Python was designed to be easy to learn: Python was born to be readable.

else if or elif

There is also the possibility of performing multiple comparisons within one if then else structure. To do this use else if or elif in python. Here is an example:

#this program tests if the number stored in  x to be greater than 3
# and if it is then set it to a maximun value of 3
x = 3.9
if  float(x)  == 3:         # the colon tells python that a code block has begun
    print('x was 3')
elif  x > 3:                      # another code block so you need another colon at the end of the line
    print('x was greater than 3')
    x = 3.0                         ## redefine x  as 3
elif x < 3:                       # another code block so you need another colon at the end of the line
    print('x was less than 3. So I letter it pass.')
else:                            # another code block so you need another colon at the end of the line
  pass                      #this command is only requred if this else or elif choise is empty

msg = 'x is now {:.1f}'.format(x)   ## some pretty printing of the result
print(msg)

x was greater than 3
x is now 3.0
  • what happens when you change what x = equal to? Try the values 4.5, -12 and 0
  • what happens when you set x = 'v' (a string)?
    • this raises an error and crashes the program
    • we will learn to deal with this later
    • Why? you can not compare a number (floating point or integer) and a string
    • Comparisons need to be made between the same types of python objects (type whos in a cell to get a list of variables and their types)

Boolean logical combinations

Another way to combine multiple conditional comparisons is to combine two comparisons with Boolean expressions or operators

Extra reading on Boole and his Logic

Some example ideas:

  • The 3 basic Boolean operators are
    • and
    • or
    • not
  • If you speak English then you know what these mean
  • and both sides of the and needs to be true for the whole statement is true
  • or either side of the or needs to be true for the whole statement is true
  • not is often called an inverter. Its presence in a statement inverts the condition
    • True becomes False
    • False becomes True
  • The same ideas apply to programming languages
  • Here is an example:
#this program checks to see if x to be greater than 3 and  less than 1
x = 5
if x <= 3 & x >= 1: #both must be true to for the  conditional action to be executed
    print('x was between 1 and 3') #conditional action
else:
    print('x was outside of the range 1 and 3')

msg = 'x is  {:.1f}'.format(x)
print(msg)

x was outside of the range 1 and 3
x is  5.0

NOTE: the syntax (symbols and commands) used in Python for Boolean operators are either the simple English words (and, or , not) or 3 symbols:

Boolean operator word for symbol symbol
and ampersand &
not bang !
or pipe or vertical bar https://en.wikipedia.org/wiki/Vertical_bar

General form

Note that the elif and else parts are optional, but make your code cooler looking.

if condition:
   code block
elif other condition:      #optonal
   code block                #optional     
else:                        #optional
  defaultcode block         # optional


Learning Review:

Write a function that

  • takes one input
  • determines if the number is greater than 0, less than zero, or zero.
  • and prints the string '+', 'nil', or '-' depending upon the conditional testing results

Learning Review: Answer

def numtest(x):
    if x < 0:
        msg = '-'
    elif x > 0:
        msg = '+'
    elif x == 0:
        msg = 'nil'
    print(msg)
    return


## test the function
numtest(75)


+

Strings, Lists

Strings

Before we jump into lists and arrays let us start with strings. We have already seen some string usage in this tutorial, but some ideas about lists can be learned from using stings.

First a string is surrounded by either single or double quotes. They both work, but you must used either single or double to define one string. Don't mix them! A string contains zero, one or more characters. For example:

word = 'python'

The variable word contains a string with 6 characters in it.

There are some special characters that can be in strings.

  • \n results in a newline
  • \t results in a tab
  • \' puts a single quote mark in the string

Here is some usage

words = 'This is a string in \t \t python. \n\t Isn\'t it awesome!'
print(words)

This is a string in 	 	 python. 
	 Isn't it awesome!

Sometimes you want to put special characters in the string as normal characters. Here is an example and a solution

  • use the r before the string definition to specify that it is raw and not interpreted.
path = 'C:\some\name_of_directory'
print(path)

betterpath = r'C:\some\name_of_directory'
print(betterpath)
C:\some
ame_of_directory
C:\some\name_of_directory

Concatenation

Strings can be concatenated.

What is Concatenation?

  • The linking of things together in a chain or series
  • Strings are concatenated with the + between strings

Here are some examples:

## un un un + ium
name = 3 * 'un' + 'ium'  # NOTICE the + between the un and ium + means concatenate in strings
print(name)
unununium

NOTE: notice the 3* 'un' in the code above. Some operators (multiply) can be used on strings. These operations mean to take 3 of the strings in a row, or concatenate 3 of the strings. Try using other operators on strings (divide, addition, subtraction), you will notice an error.

indexing and strings

From the python tutorial: https://docs.python.org/3/tutorial/introduction.html#strings

Strings can be indexed (subscripted), with the first character having index 0. (Some programming languages have a separate character type of variable. In Python there is no separate character type; a character is simply a string of size one. If you don't understand this set of statements, you can ignore it.)

Here are a bunch of examples of indexing strings:

word = 'Python'
print(word[0])  # print the character in position 0

print(word[5])  # print character in position 5

##extraction of subsets of the string
x = word[1]   #extract the second letter in the string or charactor as index = 1
print('the string\'s second letter or index 1 is',x)

x = word[-1]  # extract the last character in the string 
print('last is ',x)

x = word[-2]  # extract the second-last character in the string 
print('second to last is',x)

P
n
the string's second letter or index 1 is y
last is  n
second to last is o

NOTE: Since -0 is the same as 0, negative indices start from -1. NOTE: Indices can have positive and negative values (see the last to examples in the code block above.)

slicing stings

In addition to indexing, slicing is also supported. While indexing is used to obtain individual characters, slicing allows you to obtain a substring:

word = 'Python'
### substrings
x = word[0:2]  # characters from position 0 (included) up to 2 (excluded)
print('substring',x)

x = word[2:5]  # characters from position 2 (included) up to 5 (excluded)
print('substring',x)

print(word[:3] + word[3:])  ## see note below
substring Py
substring tho
Python

Note how the start is always included, and the end always excluded. This makes sure that s[:i] + s[i:] is always equal to s. This behavior is the main reason why python starts its indexing at zero. Here are some more demonstrations:

word = 'Python'
### substrings concatenation
x = word[:2] + word[2:]  
print('notice that ',x,' is the same as ',word)

x = word[:4] + word[4:]  
print('Also, notice that ',x,' is the same as ',word)

notice that  Python  is the same as  Python
Also, notice that  Python  is the same as  Python

Slice indices have useful defaults;

  • an omitted first index defaults to zero,
  • an omitted second index defaults to the size of the string being sliced or the end of the string.
word = 'Python'
### substring slicing
x = word[:2]   # characters from the beginning to position 2 (excluded)
print('characters from the beginning to position 2 (excluded) are => ',x)

x = word[4:]   # characters from position 4 (included) to the end
print('characters from position 4 (included) to the end are => ',x)

x = word[-2:]  # characters from the second-to-last (included) to the end
print('characters from the second to last (included) to the end are => ',x)

characters from the beginning to position 2 (excluded) are =>  Py
characters from position 4 (included) to the end are =>  on
characters from the second to last (included) to the end are =>  on

One way to remember how slices work is to think of the indices as pointing between characters, with the left edge of the first character numbered 0. Then the right edge of the last character of a string of n characters has index n, for example:

  +---+---+---+---+---+---+
  | P | y | t | h | o | n |
  +---+---+---+---+---+---+
i   0   1   2   3   4   5   6
j  -6  -5  -4  -3  -2  -1

  • The first row of numbers gives the position of the indices 0…6 in the string;
  • The second row gives the corresponding negative indices.

The slice from i to j consists of all characters between the edges labeled i and j, respectively.

Note: For non-negative indices, the length of a slice is the difference of the indices, if both are within bounds. For example, the length of word[1:3] is 2.

Attempting to use an index that is too large will result in an error. Type and see what happens:

word = 'Python'
word[42]

In contrast: out of range slice indexes are handled gracefully when used for slicing:

word = 'Python'
x = word[4:42]
print(x)
print(word[42:])    ##what is this doing prints an empty string
on

Learning Review

Using this template above: Figure out how to extract some of these examples from the word variable (word = 'Python'):

  • hon
  • Pyth
  • Pytho
  • n
  • tho
  • ytho

Learning Review: Answers

word = 'Python'
print(word[3:])
print(word[:4])
print(word[:5])
print(word[-1])
print(word[2:5])
print(word[1:5])

hon
Pyth
Pytho
n
tho
ytho

the immutable string

Python strings cannot be changed — the computer science word for this is immutable. Therefore, assigning to an indexed position in the string results in an error: Here is a definition: https://docs.python.org/3/glossary.html#term-immutable

Here is an example in which we need to lower case the first letter of Python.

word = 'Python'     #define word the first time no problem
word[0] = 'p'       # try to change it you will get an error message
word = 'python'     # redining the whole variable is no problem


Here is another way to fix this particular problem (see above).

word = 'Python'
print('p' + word[1:])


python

determine the length of a string

use the len() function to get the length of things:

word = 'supercalifragilisticexpialidocious'
length = len(word)
print('the word ', word,' has ',length,' characters.')

the word  supercalifragilisticexpialidocious  has  34  characters.

Learning Review A:

Write a function that compares the lengths of two strings and returns the longest word with its length.

  • function
  • 2 inputs
  • compute length of each string
  • compare lengths
  • output string and its length (HINT: last line of the function should include return string, length )

Learning Review B: More Advanced (to do at your own pace)

Write a function that compares the lengths of two strings and returns the longest word with its length.

  • function
  • 3 inputs
  • compute length of each string
  • compare lengths
  • output string and its length (HINT: last line should include return string, length )

lists

From Python tutorial https://docs.python.org/3/tutorial/introduction.html#lists

Python knows a number of compound data types, used to group together items. The most versatile is the list, which can be written as a list of comma-separated items (numbers, strings, other lists, whatever) between square brackets [ ] . Lists might contain items of different types, but usually for this tutorial the items all have the same type.

Defining lists

squares = [1, 4, 9, 16, 25]  ## define a list by enclosing csv in []
print(squares)
letters = ['A','B','C','D']
print(letters)
[1, 4, 9, 16, 25]
['A', 'B', 'C', 'D']

indexing and slicing lists

Just like strings in a previous section (and all other built-in sequence types), lists can be indexed and sliced:

squares = [1, 4, 9, 16, 25]  ## define a list by enclosing csv in []


x = squares[0]  # indexing returns the item
print('x =', x)

y = squares[-1]
print('y =', y)

z = squares[-3:]  # slicing returns a new list
print('z = ',z)

x = 1
y = 25
z =  [9, 16, 25]

All slice operations return a new list containing the requested elements. This means that the following slice returns a new copy of the list:

squares = [1, 4, 9, 16, 25]  ## define a list by enclosing csv in []

x = squares[:]
print('x = ',x)

x =  [1, 4, 9, 16, 25]

Concatenation of lists

Lists also support operations like concatenation:

squares = [1, 4, 9, 16, 25]  ## define a list by enclosing csv in []

x = squares + [36, 49, 64, 81, 100]  # concatentation with +
print('now x = ',x)

now x =  [1, 4, 9, 16, 25, 36, 49, 64, 81, 100]

The multiply operator on a list behaves like strings. Here is a demonstration:

squares = [1, 4, 9, 16, 25]  ## define a list by enclosing csv in []

x = 3* squares  # repeated concatentation with *
print('now x = ',x)

now x =  [1, 4, 9, 16, 25, 1, 4, 9, 16, 25, 1, 4, 9, 16, 25]

NOTE: this behavior of lists always seems weird to me. It is not what I would want to happen. We will learn later several ways to operate on all of the elements in a sequence (List, arrays, dataframes).

Lists are Mutable

Unlike strings, which are immutable, lists are a mutable type, i.e. it is possible to change their content:

squares = [1, 4, 9, 16, 25]  ## define a list by enclosing csv in []

cubes = [1, 8, 27, 65, 125]  # something's wrong here
print('something is wrong here:  ',cubes)
print('but ',4 ** 3, ' is the cube of 4')  # the cube of 4 is 64, not 65!

cubes[3] = 64  # replace the wrong value



print('now cubes  = ',cubes)

something is wrong here:   [1, 8, 27, 65, 125]
but  64  is the cube of 4
now cubes  =  [1, 8, 27, 64, 125]

Adding items to the end of a list

You can also add new items at the end of the list, by using the append() method (we will see more about methods later)

squares = [1, 4, 9, 16, 25]  ## define a list by enclosing csv in []

cubes = [1, 8, 27, 65, 125]  # something's wrong here
print('cubes was initially defined as ',cubes)

cubes.append(216)  # add the cube of 6
cubes.append(7 ** 3)  # and the cube of 7
print('appending 2 new values to cubes.  now cubes is ',cubes)


cubes was initially defined as  [1, 8, 27, 65, 125]
appending 2 new values to cubes.  now cubes is  [1, 8, 27, 65, 125, 216, 343]

Assignment to slices is also possible, and this can even change the size of the list or clear it entirely:

letters = ['a', 'b', 'c', 'd', 'e', 'f', 'g']
print('letters is ',letters)

# replace some values
letters[2:5] = ['C', 'D', 'E']
print('replacing indices 2-5 yields',letters)

# now remove them
letters[2:5] = []
print('droping the indeces 2-5 yeilds',letters)


# clear the list by replacing all the elements with an empty list
letters[:] = []
print('clearing the whole list yeilds and empty list',letters)
letters is  ['a', 'b', 'c', 'd', 'e', 'f', 'g']
replacing indices 2-5 yields ['a', 'b', 'C', 'D', 'E', 'f', 'g']
droping the indeces 2-5 yeilds ['a', 'b', 'f', 'g']
clearing the whole list yeilds and empty list []

Getting the length of a list

The built-in function len() also applies to lists:

letters = ['a', 'b', 'c', 'd']
length = len(letters)
print('The length of this list is ',length)
The length of this list is  4

Lists of Lists

It is possible to nest lists (create lists containing other lists), for example:

### hand coding a list of lists

LL = [['I','think', 'we'],['are','here']]
print('hand coded list of lists ',LL)

## or more programatically
a = ['a', 'b', 'c']    #make 1 list
n = [1, 2, 3]          #make another list
x = [a, n]             # now make a list of lists   or nested lists
print('look at what x is ',x)

##indexing a nested list
b = x[0]
print('the 0th index in x is ',b, ' and is the sublist a =>',a)

## indexing a nested list
deepinside = x[0][1]  ## note the 2 sets of square brackets 
print('here is an item within the 0 index nested list = >',deepinside)
print('here is an item within the 1 index nested list = >',x[1][2])

hand coded list of lists  [['I', 'think', 'we'], ['are', 'here']]
look at what x is  [['a', 'b', 'c'], [1, 2, 3]]
the 0th index in x is  ['a', 'b', 'c']  and is the sublist a => ['a', 'b', 'c']
here is an item within the 0 index nested list = > b
here is an item within the 1 index nested list = > 3

learning Review

Create a funny sentence function. This function will pick 3 random indices and pull words from 3 random lists, then print to the screen the concatenated funny sentence.

Example: list of subjects: I, you, donkey list of verbs: eat, like, poop list of objects: bread, corn, hippos

randomly choose three indices 0,1,2 then your funny sentence that should be printed to the screen is:

I like hippos

Use the numpy function randint to get an integer that can be used as an index.

import numpy as np                 #this command loads all of the functions in numpy and labels them np

from numpy.random import randint  ## this command loads the function randint in to memory for later use

subjects = ['I', 'You']
idx = randint(0,high=2)       # the 0 is the lowest integer (0th index)
                                  # ths high=3 is one count above the highest integer that you need
                                  # to make this flexible maybe don't hard code the high = 2
                                  # but get it from the list     

print('idx is ',idx)

idx is  1

Assignment 6 (ass6) Part 1: create a function that

  • takes 2 inputs
    • 1\(^{st}\) is a list of floating point numbers (1.0, 3.12, etc.)
    • 2\(^{nd}\) is an index that is less than the length of the list
  • takes the number in the list at the index position and replaces it with 13
  • appends the number 114 to the end of the list
  • uses the built-in function sum() to sum all of the items in the list
  • returns the sum of the newlist

Due Date is in the directory

Assignment 6 (ass6): Part 2 List slicing (10 points)

create a function that

  • takes 1 input
  • the input is a list of items 'dog',1,'R',1.0, 3.12, or whatever with length > 5
  • checks to see if the list is > 5 items long
    • prints an error message if the input list is too short, and returns a 1
  • otherwise takes a slice of the 1,2,4 index of the list
  • returns the returns the slice of the list

Due Date is in the directory

control structures: for loop

This control structure is a little difficult to compare to a children's game. The most similar game is a relay race. Your goal in this race is to carry a stick around the track a specific number of times.

for each time around the track
carry the baton
when you have looped around the track the required amount of times the race is at an end

Use a for loop to do a task a specific number to times. Here is an example

for i in [1,2,3,4,5,6,7,8,9,10]:
    print(i)
1
2
3
4
5
6
7
8
9
10

In the previous example, [1,2,3,4,5,6,7,8,9,10] is a list. I contains a list of integer numbers.

You would read this code block as "for i equals items in the list", [1,2,3,4,5,6,7,8,9,10], then display i then end the code block. In a for loop, the iterator variable i in the case above takes on the values from the list 1, 2, 3, 4, 5, ….

range()

What would be really nice is if we had a way to generating this list [1,2,3,4,5,6,7,8,9,10] rather than typing it. Use the command range()

This example should give the same output.

for i in range(10):
    print(i)
0
1
2
3
4
5
6
7
8
9

The range() function can take up to 3 inputs.

  • range(begin,end,stepsize)

Let's use this range() function to add up some numbers

start = 1
stop = 101
step = 1
count = 0                           ## this is the seed of the count
for i in range(start,stop,step):
    count =  count + i
    print(i, count)

1 1
2 3
3 6
4 10
5 15
6 21
7 28
8 36
9 45
10 55
11 66
12 78
13 91
14 105
15 120
16 136
17 153
18 171
19 190
20 210
21 231
22 253
23 276
24 300
25 325
26 351
27 378
28 406
29 435
30 465
31 496
32 528
33 561
34 595
35 630
36 666
37 703
38 741
39 780
40 820
41 861
42 903
43 946
44 990
45 1035
46 1081
47 1128
48 1176
49 1225
50 1275
51 1326
52 1378
53 1431
54 1485
55 1540
56 1596
57 1653
58 1711
59 1770
60 1830
61 1891
62 1953
63 2016
64 2080
65 2145
66 2211
67 2278
68 2346
69 2415
70 2485
71 2556
72 2628
73 2701
74 2775
75 2850
76 2926
77 3003
78 3081
79 3160
80 3240
81 3321
82 3403
83 3486
84 3570
85 3655
86 3741
87 3828
88 3916
89 4005
90 4095
91 4186
92 4278
93 4371
94 4465
95 4560
96 4656
97 4753
98 4851
99 4950
100 5050

Here is another way to do the same thing, but it stores the values of i in a list:

start = 1
stop = 101
step = 1
L = []                         ## this starts a list to put the numbers in
for i in range(start,stop,step):
    L.append(i)
print(i, L)
Ls = sum(L) 
print('The sum is ', Ls)

100 [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100]
The sum is  5050

NOTE: the command sum(). It does what you think it should. It sums the list L.

List processing with for loops with range()

This list generator (range()) is good for generating indices for accessing components in a list or other sequence data structure.

L = ['The','cat', 'has','a','scary','smell']
for i in range(len(L)):
  print(L[i])

The
cat
has
a
scary
smell

In Python, this iterating over a list is very powerful, especially because the list can contain any python structure.

  L = ['The','cat', 'has','a','scary','smell']
  for word in L:
    print(word)
#Notice how there is no math here at all


We can also access more than one list by using an indices generator.

L = ['The','cat', 'has','a','scary','smell']
P = ['THE','CAT', 'HAS','A','SCARY','SMELL']
for i,word in enumerate(L):  ##use the enumerate around a list to create an index generator
    print(i,word,P[i])       # i is the index of the list and word is the item in the list


0 The THE
1 cat CAT
2 has HAS
3 a A
4 scary SCARY
5 smell SMELL

Learning Review: add up some stuff

Create a function to add up all of the numbers in a range

  • use a for loop to step through the range of numbers
  • use range: 12 to 115
  • also figure out a quick change to your range of every 3rd number
  • print and returns the sum

Learning Review: Answer

# learning review check add up some stuff
#first do it not in a function
start = 12
stop = 115 
step = 1
count = 0
for i in range(start, stop + 1, step):
    count = count + i
print(i,count)

## then make a function
def addmup(start,stop, step):  #put the parts defined by the user as inputs
    '''
    Adds up all the numbers from start to stop 
    step is the step size between the start and stop, 1,2,3, etc.
    '''
    count = 0                 ## you start adding at zero
    for i in range(start, stop + 1, step):  ## notice stop + 1, range before the value in stop
        count = count + i       
    print('The sum of all values in the range (',start,' to ',stop,') is',count)
    return count

## test out your function
addmup(12,115,1)


115 6604
The sum of all values in the range ( 12  to  115 ) is 6604

Learning Review: Fibonacci list

write a function that

  • takes 1 input as an integer (n)
  • creates a list that contains the Fibonacci sequence up to and including the nth
  • returns the list

Learning Review: Fibonacci list: Answer

def fibonacci(n):
    '''
    function to generate a list of the first nth numbers in the Fibonacci sequence
    n is the integer number in the fibonacci sequence
    n must be larger than 1
    '''

    F1 = 1              #first number in the Fibonacci sequence
    F2 = 1              # second number in the Fibonacci sequence
    fiblist = [F1,F2]        # seed the list with the first two fibonacci numbers
    for i in range(1,n-1):      # start loop up to n, note:  start with 1 because that is how most people count, n-1 already have fisrt 2 in list, remember in range(stop) the stop number is 1< stop
        F = F1+ F2       # calculate the next number
        fiblist.append(F)  #store the F in running list
        F1= F2              # Shift current F1 to F2
        F2 = F              # shift current F to F1   
    return fiblist          #return the list 

print(fibonacci(12))

[1, 1, 2, 3, 5, 8, 13, 21, 34, 55, 89, 144]

Check if an item is in a list

if 'cat' in ['cat','dog','cow']:
    print('yes')



yes

ass7: Fibonacci 1

Write a function that will

  • take 1 input that is a positive integer (n)
    • if the input is not an integer return an error of 1 (HINT: use the isinstance function link)
    • and print an error message
  • calculate the nth Fibonacci number
  • return the Fibonacci number

ass7: Fibonacci 2

write a function that

  • takes an integer as an input
  • checks if the input is a Fibonacci number
  • returns a True if it is a Fibonacci number
  • returns a False if it is not a Fibonacci number

ass7: Fibonacci 3

Format in Markdown in the cell provided the Equation:

\begin{equation} F_n = F_{n-1} + F_{n-2} \end{equation}
  • Due Date Feb 16, 2018 at noon

Arrays and Matrices for data storage

As we learned in the last several modules, lists are powerful Python objects that can be used to store information or data. However there are several (2) limitations with lists that make them less useful for storing data.

Specifically,

  • mathematical operations on a list don't work in a useful way.
  • ???

Here is an example this limitation: Let's say you were needed to calculate the square of a list of numbers (such as in the calculation of the standard deviation.)

Try this in python

data = [15.2, 16.7, 17.34, 14.6, 13.9]
data*data   # this is the calculation for the square of a number
data ** 2  # nor can you do this

Notice how it gives an error. You can not multiply 2 lists together nor can you take the power 2 of a list.

Here are 2 ways of doing this operation using a for loop Note there are more variables used here than are necessary. I am using them to make the example more understandable.

data = [15.2, 16.7, 17.3, 14.6, 13.9]   ## the org data
sqdata = [ ]                            ## create an empty list to store the squares
for d in data:                          ## us 'in' to get the items out of the data list one at a time
    sqd = d**2                          ## square the data and store it in a new variable
    sqdata.append(sqd)                  ## append the square to the list 
print(data)
print(sqdata)

[15.2, 16.7, 17.3, 14.6, 13.9]
[231.04, 278.89, 299.29, 213.16, 193.21]
data = [15.2, 16.7, 17.3, 14.6, 13.9]       ## the org data
sqdata = [ ]                                ## create an empty list to store the squares
for i in range(len(data)):                  ## create an increasing index 'i' to access the items in data
    d = data[i]                             ## extract the item from list 'data'
    sqd = d*d                               ## square the data and store it in a new variable
    sqdata.append(sqd)                      ## append the square to the list 

print(data)
print(sqdata)

[15.2, 16.7, 17.3, 14.6, 13.9]
[231.04, 278.89, 299.29, 213.16, 193.21]

Using a for loop to process lists can be tedious. If you have a huge list it is also very slow. (This is a fact that is not self-evident from what we have covered so far. You will have to either accept this statement or you can read about it link.

So the Python community have created additional Python object types to help with this kind of problem. They are array, matrix, and dataframe. These are all very similar to List, but they have special properties that are useful for data analysis.

You might envision arrays as a kind of data table or spreadsheet-type of computer object.

Look at this table.

1
2
3
4

It is basically an array of dimensions 4 by 1 (4x1) or (4,1). [ Notice the dimensions are (row, column). ] And here

1 2 3 4

is an array of dimensions 1 by 4 (1x4) or (4,1). These are both considered 1D arrays, which means that they only have 1 dimension that is greater than 1.

Now consider the following array

1 2
3 4

What are its dimensions?

  • ANSWER: (2,2) two by two

The only limitation with thinking of arrays as a kind of table is that most people don't think of data tables as having more than 2 dimensions. But arrays can have as many dimensions as you can imagine (or whatever the limit of the computer language Creator's imagination was).

Here is some code for understanding arrays.

creating arrays

To create/operate/use arrays we must use the numpy module.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions

import numpy as np                 #this command loads all of the functions in numpy and labels them np


A = np.array([4,5,6,7])            #turn a list into an array
print(type(A))                     ## print the type of object 


print('the shape of the array is ',A.shape)               # print the shape of the array.
<class 'numpy.ndarray'>
the shape of the array is  (4,)

Type %whos in the cell of Jupyter (the % means magic command in Jupyter) and you should see in the list:

Variable   Type       Data/Info
-------------------------------
A          ndarray    4: 4 elems, type `int64`, 32 bytes
np         module     <module 'numpy' from '/ho<...>kages/numpy/__init__.py'>

Notice the variable A type is ndarray. You can also access this information with the type(function). ndarray means n-dimensional array.

Use the .shape method to extract the shape of an array. Why does it say (4,) instead of (4,1)? This is just Python's way of saving space/memory.

Element-wise Operations on an array

The reason of using arrays is to operate on them easily. Let's take the problem of squaring a list above. Here is how to solve this problem with numpy arrays.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions

import numpy as np                 #this command loads all of the functions in numpy and labels them np
datalist = [15.2, 16.7, 17.3, 14.6, 13.9]       ## the org data
data = np.array(datalist)                     ## first make the list into an array 
sqdata = data**2                 ## square the data
print(datalist)
print(sqdata)
print(type(sqdata))
[15.2, 16.7, 17.3, 14.6, 13.9]
[231.04 278.89 299.29 213.16 193.21]
<class 'numpy.ndarray'>

Whenever you use an operator on an np.array object it operates on each element and returns a new array. Try a several

  • data-sqdata
  • data*sqdata
  • data + data

These all do element-wise calculations.

indexing and slicing

Array indexing and slicing works just like a list.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions

import numpy as np                 #this command loads all of the functions in numpy and labels them np

data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] )  # onestep create of array

print('The zeroth index of the array is ', data[0])    #indexing returns a number
print('The first 3 items in the array are ', data[:3])  ## slicing returns an array

The zeroth index of the array is  15.2
The first 3 items in the array are  [15.2 16.7 17.3]

modification of an array

*Array*s are like *list*s in that they can be modified.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions

import numpy as np                 #this command loads all of the functions in numpy and labels them np

data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] )  # onestep create of array

print('The zeroth index of the array is ', data[0])    #indexing returns a number

data[0] = 55.5                     ## replace this zeroth indexed number

print('The zeroth index of the array is now', data[0])    #indexing returns a number


The zeroth index of the array is  15.2
The zeroth index of the array is now 55.5

Special arrays

There are some special arrays that are easy to create with numpy.

  • an array of all ones
  • an array of all zeros
  • an array of all random numbers
# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions

import numpy as np                 #this command loads all of the functions in numpy and labels them np

## all ones
A = np.ones( (3,3) )  ## the (3,3) is the shape that you want

print('A is \n', A)  ## the \n prints a new line
print(A.shape)

## all zeros
B = np.ones( (4,4) )  ## the (4,4) is the shape

print('B is \n', B)    ## the \n prints a new line
print(B.shape)

## normally distributed random numbers
center = 3       ##center of random numbers
spread = 2       ## standard deivation of the random numbers 
shape = (2,4)    ## shape of the array created.  this is called a tuple
C = np.random.normal(loc=center,scale=spread,size=shape)
print('C is \n',C)  ## the \n prints a new line
print(C.shape)

## uniformly distributed random numbers
D = np.random.rand(3,5)                ## random numbers are spread out evenly between 0 and 1
print('D is \n ', D)         ## the \n prints a new line
print(D.shape)

A is 
 [[1. 1. 1.]
 [1. 1. 1.]
 [1. 1. 1.]]
(3, 3)
B is 
 [[1. 1. 1. 1.]
 [1. 1. 1. 1.]
 [1. 1. 1. 1.]
 [1. 1. 1. 1.]]
(4, 4)
C is 
 [[2.73534909 3.65547042 0.63413359 0.98220572]
 [0.74696185 0.11992447 1.89381477 3.80934037]]
(2, 4)
D is 
  [[0.86655031 0.5470224  0.66410465 0.57506126 0.35278717]
 [0.31438514 0.77715677 0.89651335 0.89601613 0.88361964]
 [0.13078894 0.40011216 0.58053045 0.14277761 0.59582208]]
(3, 5)

Whole array manipulations (continued)

Another powerful utility of numpy is its built-in functions. Here are a few useful examples that do what you think they do.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions

import numpy as np                 #this command loads all of the functions in numpy and labels them np


data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] )  # onestep create of array
print(data)
print(np.sin( data ))  
print(np.sum(data))
print(np.median(data))
print(np.min(data))         ## determines the miniumn value in an array
print(np.max(data))         ## determines the maximun value in an array
print(np.std(data,ddof=1))  ## determine the standard deviation "ddof = 1"  means use degree of freedom = 1, which is the words for n-1 in the bottom of the std. dev. equation.


[15.2 16.7 17.3 14.6 13.9]
[ 0.48639869 -0.83714178 -0.99977443  0.89479117  0.9720075 ]
77.7
15.2
13.9
17.3
1.4258330898110059

numpy has about 99% of all of the functions that you might want. Here is a list in their reference page.

array methods

In the numpy module, the array objects have methods associated with them. A method is an operation on an object that can be applied to that object. Jupyter has several nice features that can help use understand some of these.

import numpy as np  #this command loads all of the functions in numpy and labels them np

data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] )  # onestep create of array

print( np.sum(data) ) ## print the np.sum of the data ndarray

print( data.sum() )  ## same as np.sum( )  but as a method: the dot .sum() makes it a method




77.7
77.7

The methods produce the same results as the functions. Note: Not all functions are methods. Try these methods:

  • sum()
  • mean()
  • max()
  • sin()

In Jupyter there are some convenience features. Type

import numpy as np  #this command loads all of the functions in numpy and labels them np

data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] )  # onestep create of array

# then type data. and depress the tab button, and wait,  you will see a menu pop-up of all the methods

the Matrix

Python provides a special type of array called a matrix. The major differences between a matrix and an array is the manner in which mathematical operations are handled.

  • Arrays review: math works element-wise
    # many functions are avaible in modules or libraries
    # in this example we will load the numpy module of functions
    
    import numpy as np                 #this command loads all of the functions in numpy and labels them np
    
    A = np.array([[ 2,  2.], [-2. , 2]] )    ## create array
    print(A*A)                                ## perform element-wise multiplication
    
    B = np.array([2,2])
    print('The shape of B is', B.shape)        ## notice how the shape is (2,) 
    
    #you can also flip the rows and columns, this process is called transpose
    print('B is \n', B.T)    # .T method for transpose
    print(B.T.shape)   ## notice the shape now,  transpose does not mean anything for a (2,) array
    
    
    [[4. 4.]
     [4. 4.]]
    The shape of B is (2,)
    B is 
     [2 2]
    (2,)
    
  • Matrix: math works using the rules of Linear Algebra

    We will learn some basics of linear algebra later in this tutorial, but for now here is an example of this linear algebra behavior of matrices. Unusual results from mathematical operation can result.

    # many functions are avaible in modules or libraries
    # in this example we will load the numpy module of functions
    
    import numpy as np                 #this command loads all of the functions in numpy and labels them np
    
    A = np.matrix([[ 2,  2.], [-2. , 2]] )    ## create a matrix
    print(A*A)                                ## perform matrix multiplication, not element-wise multiplication
    
    B = np.matrix([2,2])
    print('The shape of B is', B.shape)                            ## notice how the shape is (2,1) and not (2,) like an array
    
    #you can also flip the rows and columns, this process is called transpose
    print('B is \n', B.T)    # .T method for transpose
    print(B.T.shape)   ## notice the shape now
    
    
    # you can also convert between the array and the matrix data objects
    Am = np.matrix([[ 2,  2.], [-2. , 2]] )
    print('Matrix type', type(Am))
    Aa = np.array(Am)
    print('Now it is an array', type(Aa))
    
    # the reverse is also true
    Aa = np.array([15.2, 16.7, 17.34, 14.6, 13.9])
    print('Array data type', Aa)
    Am = np.matrix(Aa)
    print('Now it is a matrix', type(Am))
    
    [[ 0.  8.]
     [-8.  0.]]
    The shape of B is (1, 2)
    B is 
     [[2]
     [2]]
    (2, 1)
    Matrix type <class 'numpy.matrix'>
    Now it is an array <class 'numpy.ndarray'>
    Array data type [15.2  16.7  17.34 14.6  13.9 ]
    Now it is a matrix <class 'numpy.matrix'>
    

    Some matrix specific operations are diag and trace. Here are some examples:

    # many functions are avaible in modules or libraries
    # in this example we will load the numpy module of functions
    
    import numpy as np                 #this command loads all of the functions in numpy and labels them np
    
    A = np.matrix([[ 2,  2.], [-2. , 2]] )    ## create a matrix
    print(A)
    
    
    
    print( np.diag(A) )                    ## the diagonal of the matrix is an array
    D = np.diag(A)
    print( np.sum(D) )                  ## sum of the diagonal is the trace
    
    # or
    
    T = np.trace(A)
    print(T)
    
    
    [[ 2.  2.]
     [-2.  2.]]
    [2. 2.]
    4.0
    4.0
    

Assignment 8 Arrays (ass8)

create a function that will

  • take as inputs 2 arrays of the same size
  • element-wise add the two arrays together
  • divide resulting sum-array by 2
  • return the array return the resulting array

Assignment 9 Matrix math (ass9)

create a function that will

  • take as inputs 2 matrices of the same size
  • matrix multiply them together
  • then calculates the trace for product of the two matrices
  • returns the trace of the matrix multiplication (or the sum of the diagonal)

Dataframes (the named arrays)

Dictionaries adfad

Dictionaries are another kind of information/data storage object in python. We will not use them very much in this tutorial. They are a good segue into the next section dataframes, so that is why we are covering them here.

Python dictionaries are just like dictionaries in the real world. You put in a word and you get out a definition.

A dictionary is defined by either the word dict() or enclosing it in {}. Here are some examples.

Just like a list you can store anything in a dictionary.

import numpy as np  #this command loads all of the functions in numpy and labels them np

## generic example:
#      name = {key:value,key:value}

hr = {'Aaron':47,
      'Ruth':60,
      'Bonds':73,
      'McGwire':70,
      'Sosa':66
     }

print( hr['Bonds'] )   

### a simplified spectroscopic example
wavelengths = np.array([200, 250, 300, 350, 400])
spec = np.array([0, 0.2,     0.5, 0.3,  0.001])

d = {'x':wavelengths, 'y':spec}

y = d['y']
print(y)

73
[0.    0.2   0.5   0.3   0.001]

Dataframes

Python has another module called Pandas user guide. That was built to make certain tasks in python/numpy easier. In this module is defined a new data type object called a dataframe. You can think of these types as named arrays, or some kind of combination of array and dictionary.

Here is a continuation of the baseball example above:

import numpy as np  #this command loads all of the functions in numpy and labels them np
import pandas as pd   ## this command loads the pandas module

hr = {'Aaron':47,
      'Ruth':60,
      'Bonds':73,
      'McGwire':70,
      'Sosa':66
     }

df = pd.DataFrame.from_dict(hr,orient='index',columns=['HR'])

print(df)  #the dataframe presents itself as a table with the rows labeled (called the index) and the column(s) labeled
print('')
print('****************** spacer *********************')
print('Average Homeruns \n',df.mean())  ## you can operate on the dataframe, NOTE:  the output of these operations is another dataframe


print('****************** spacer *********************')
#to get the values out of the dataframe use the .values method


hrdf = df.mean()
print(hrdf)
print(hrdf.values)  ## output as an array


         HR
Aaron    47
Ruth     60
Bonds    73
McGwire  70
Sosa     66

****************** spacer *********************
Average Homeruns 
 HR    63.2
dtype: float64
****************** spacer *********************
HR    63.2
dtype: float64
[63.2]

Here is another example that is a little more relevant to chemistry

import numpy as np  #this command loads all of the functions in numpy and labels them np
import pandas as pd   ## this command loads the pandas module


wavelengths = np.array([200, 250, 300, 350, 400])  ## wavelength labels for absorbance values
spec1 = np.array([0, 0.2,     0.5, 0.3,  0.001])  ## absorbance values
spec2 = np.array([0.2,0.5,0.1,0.0,0.0])           ## aborbance values


d = {'spectrum 1':spec1,
    'spectrum 2':spec2}                      ## make a dictionary

df = pd.DataFrame.from_dict(d,orient='index',columns=wavelengths)   ## make a dataframe

print('Type for spec',type(spec1))
print('Type for d',type(d))

print('Type for df',type(df))
print('****************** spacer *********************')
## as a dataframe you can take the average spectrum

print(df.mean())

print('****************** spacer *********************')
## or you can get the absorbacne values at a specific wavelength
print(df[300])
Type for spec <class 'numpy.ndarray'>
Type for d <class 'dict'>
Type for df <class 'pandas.core.frame.DataFrame'>
****************** spacer *********************
200    0.1000
250    0.3500
300    0.3000
350    0.1500
400    0.0005
dtype: float64
****************** spacer *********************
spectrum 1    0.5
spectrum 2    0.1
Name: 300, dtype: float64

CANCELED Activity Assignment 10 (ass10) (maybe move to loading data section)

Making Graphics and Figures

All discussion of spectroscopic (and most other) data is often accompanied by some kind of graphical representation of the data. This section will describe the fundamentals of graphical representation of data plus the graphical representation of the meaning of the data. \footnote{I don't actually know what this last clause means, but it sounds cool.}

It is my philosophy that all graphical data presentations must stand by themselves, without oral explanation. So when you hand a printout of data to your colleague or boss, you should not have to explain every detail of the graphic. Some obvious parts of the presentation such as axes labels, legends, and captions easily enhance the understanding of the graphic. Herein are a set of examples from which you can learn.

There are several types of data presentations, and the choice of which to use depends upon the type of data and the intended usage of the data presentation.

There are 6 basic types of data presentation graphs that are commonly used by chemists: line graph, scatter plot, histograms, contour plots, surface plots, and images. (I guess there is also a bar graph and pie graph, but I don't really know when to use these types.) There are many perturbations within each of these graph types including combinations of each within one presentation.

There are several intended uses for these data presentations, and these uses in part dictate how the figures are created.

  • Computer-based inspection: For me this is the most useful, because I can interact with the data (zoom in on specific parts, inspect the peak max or min, overlay 2 spectra, etc.).
  • Paper-based lab notebook: This printed version does not require super high resolution; I will refer to this type of presentation as a thumbnail print.
  • Publication Quality: Related to the thumbnail print is the publication quality print, which is a higher resolution file size and may also have additional annotations and figure captions.
  • Beamer Presentation: Also related to the printed version is for inclusion into a beamer presentation or image projection system for giving talks, which usually don't require a high resolution graphic format.

All of these can be created with python, numpy, matplotlib, pandas, and scipy. Additionally, these data presentations may need to be annotated to enhance the understanding of the information presented.

In the following subsections are examples of these different types of graphics and their uses.

General Philosophy for what type of graph to use

In the introduction above were listed 6 types of graphs: line graph, scatter plot, histograms, contour plots, surface plots, and images. So a question might arise in your mind, "Which type do I use?" There have been many discussions over the years about this question. Here are a few links:

This is an unsettled issue that is full of opinions. Herein I will present my opinions as well as some rules that are settled by conventions within the field of chemistry (mostly IUPAC). I will delineate between my opinions and conventions, by stating when something is my opinion. All other statements should be assumed to be based upon best practices (conventions).

These next 6 subsections are all my opinion, but I believe that most scientist would agree with them.

when to use a scatter plot

Whenever you need to fit a smallish set of data with a mathematical function, it is better to plot the data with markers, and the fit with a line. My reasons are that this formula produces a clear distinction between the actual data and the mathematical function. You don't want the viewers of you data presentation to think that you have more data than you do.

Some examples are shown in the Figure below.

multi_examplescatter.png

Figure 12: Example scatter plots with the data plotted with markers and the fit with a line

When to use a line plot

Whenever you have loads of data points that are follow a predictable trajectory, like in a spectrum or chromatorgram, using markers will make the line so thick that you might inadvertently hide the curvature. In these situations, I believe you can use a smoothly curving line to present the spectrum, chromatorgram, etc. Some examples are shown in the figure below.

multi_examplesline.png

Figure 13: Example line plots with lots of points

When to use histograms

Whenever you have 1D data that is from a set of repeated measurement, and your goal is to represent the spread (how repeatable) in the data, then a good choice is a histogram. An example, might be if you were monitoring the arsenic concentration level in bunch of samples.

multi_examplehist.png

Figure 14: Example of a histogram of repeated measurements

When to use contour and surface plots and images

In chemistry data, these 3 types of plots can be used interchangeably. One exception is when presenting data that is in fact an image (e.g. photograph, SEM image, etc.). The types of data that utilize these types of graphs fall in to 2 categories:

  • correlated/anticorrelated data

    These types of data are commonly found in 2D correlation spectroscopy. Some of you may have learned about 2D NMR measurements such as COSY ( link ) and HETCOR. These data are often presented as contour plots. These NMR measurements are a subset of experiments that are part of the generalized 2D spectroscopy type of measurements. Here is a very good article link.

    In all of these cases, the X and Y dimensions are some kind of spectrum (X and Y can be the same or different spectra or even spectral types), and the in the middle are the relative heights of the correlations between the two (X and Y) spectra portrayed as contours.

    Here is an example of a contour plot of the correlation of UV spectra during a chemical reaction.

    2dcorspec.png

    Figure 15: Example of generalized 2D spectroscopy

    This same kind of data could also be displayed as a 3D surface. In fact a contour plot is just a flattened surface plot. Here are the same data shown in 3 different graphs. Notice how the contour graph and image really need a color bar to the right so that the reader can know what the colors represent.

    multi_examplesurf_contour_image.png

    Figure 16: Examples of a contour, surface plot and an image of the same data

  • data in which multiple two axes are the same kind of data (spatial) and the 3rd is a different kind of data (concentration, spectra band height, etc.)

    You have seen this kind of data before in the following figure.

    figureprint.png

    Figure 17: Finger print, DNA, RDX

Generic Example

Each of the examples within this section of the tutorial will contain several parts. Not all of the parts will be used in every example, but the part is listed to keep the examples consistent. Here is a generic list of the parts:

1.  import modules  (you only need this part once per jupyter notebook)

2  get or make the data

3.  process the data

4. plot the data

5. plot the data

6.  save the data in the required format

7. house keeping (this is part of the tutorial but not for you to copy)

plotting useless data

Let us start by just learning to plot using made-up, simulated data. There are 3 kinds of graphs that are common in chemistry:

  • a continuous graph of data without markers
  • an xy scatter graph with markers
  • a histogram of data from repeated measurements with some spread in the data

simulated spectrum

For spectra, chromatograms, and other continuous data it is better to plot the data with a line without markers. The main reason is that the markers would be so close together that they would interfere with the viewer's perspective of the shape of the curve. The default for the matplotlib library is to add space to either sided of these types of graphs. Personally, I don't like this. For now we will leave then there, and learn how to eliminate them latter in this tutorial.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
# you only need to load the modules once per file
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module 
import matplotlib.pyplot as plt    # plotting module


## get/make the data
x  = np.arange(200,400)    # these will be the x values in the graph for a simulated  UV spectrum
xc = 250                   # this is the center of the band
sigma = 10                 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the lambda max of the curve

spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve (idea shape of lots of data)

## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4))   ## figsize is the aspect ratio of the figure
                                         ## ax is the variable for the axis object where you want to work   
ax.plot(x,spec)                          ## .plot(x,y) method to make a plot on the ax object

ax.set_xlim((200,400))                   ## set x dim edges of the plot
ax.set_xlabel('Wavelength /nm')          ## LaTeXworks here
ax.set_ylabel('Absorbance')              ## and here
ofile = 'figures/examplespec.png'        ## name of the output file
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resolution thing

### This part is for making this website, so you don't need it
return ofile

examplespec.png

simulated Beer's law data

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
C = [1.0, 2.0, 3.0, 4.0, 5.0]          #  these will be the x values in the graph they are simulating concentration in ppm
A = [0.1, 0.2, 0.3, 0.4, 0.5]    #this is the simulated Aborbance values
## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure
ax.scatter(C,A, s=55,marker='.')   ## ax is the variable for the axis where you want to work
                                   ## s is the size of the marker
                                   ## marker is the marker some choices: ('.','o', 'v', '^', '<', '>', '8', 's', 'p', '*', 'h', 'H', 'D', 'd', 'P', 'X')

ax.set_xlabel('Concentration /ppm')      ## latex works here
ax.set_ylabel('Absorbance')              ## and here
ofile = 'figures/examplescatter.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing

### This part is for making this website, so you don't need it
return ofile

examplescatter.png

Simulated Mass Data

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
mass = 10+np.random.randn(200)*0.29    ## these will be 200 simulated mass measurements normally distributed around 10 g

## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4))              ## figsize is the aspect ratio of the figure
ax.hist(mass)                                     ## ax is the variable for the axis where you want to work

ax.set_xlabel('Mass /g')               ## latex works here
ax.set_ylabel('Count')
ofile = 'figures/examplehist.png'   ## the output filename
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resolution thing

### This part is for making this website, so you don't need it
return ofile

examplehist.png

First some data

This part of computer programming is arguably the most difficult to learn for new coders. Some people call it data wrangling. These words mean getting the data from your instrument's output into the memory of python for processing. Because of this difficulty, in this course we will not deal with it. If you need help getting your data from your research into python for processing ask me, and we can do it outside of class. I am happy to assist you. I am making this teaching choice, because every instrument has its own kind of data file format, and to cover them all would take the entire semester.

So for this class, I will issue to you data in 3 kinds of formats:

  • One will be single datafiles *.csv or *.dat with or without headers stored as text files. Some examples of types of data:
    • chromatograms
    • spectra
    • concentration data
  • A relative of the *.csv file is the MS Excel files.

It is pretty common for data/results to be stored in an Excel file. I believe this is because it is pretty easy to type data into spreadsheet programs such as MS Excel. I also believe that is why they are so popular. So we will learn how to read data from these file.

  • The other type will be datasets in either *.mat files or *.hdf5 files. Some examples:
    • hyphenated technique data (GC-MS, Raman spectroscopic imaging, 2D NMR data)
    • lots of similar data (100's of IR spectra)
    • This type of file is a kind of archive in which the data has been organized for you
    • we will see some examples below

I will assume that all data is stored in a directory under your working directory named data. So all data file paths will look like

filename = 'data/somedatafilename.dat'   # on mac or unix
filename = 'data\fomedatafilename.dat'   # on a windows computer  

loading *.csv or *.dat files with no header and plot it

These will be files that contain 2 or more columns of data separated by a comma or tab or space. Traditionally, the first column is the independent values (you plot them on the x axis), and the 2nd, 3rd, etc. column(s) is(are) the dependent values (you plot them on the y-axis).

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
ifile = 'data/sample1.csv'           ## the name and relative path of the data file
rawdata = np.genfromtxt(ifile,delimiter=",")         ## read the file output into an array

print('the shape of the d is ',rawdata.shape)    ## inspect the data size

x = rawdata[:,0]                                    ## extract/slice the wavenumbers (x values)
spec = rawdata[:,1]                                 ## extract/slice the absorbance values (y-values)

## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4))    ## figsize is the aspect ratio of the figure
ax.plot(x,spec)                           ## ax is the variable for the axis where you want to work


ax.set_xlim((650,4000))            ## set graph's limits
ax.invert_xaxis()                  ## vibrational spectra need this

ax.set_xlabel('Wavenumber /cm$^{-1}$')      ## latex works here
ax.set_ylabel('Absorbance')                 ## latex works here
fig.tight_layout()                   ## removes a bunch of white space around figure
ofile = 'figures/dataspec.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing

### This part is for making this website, so you don't need it
return ofile

dataspec.png

Loading *.mat files

This file format is matlab's high performance data format. It is based upon the HDF5 format. http://www.hdfgroup.org/HDF5/

This is the faster and easier of the file formats whenever you have lots of data to read into your program. The data is already stored in the variable. To use them, load this data.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

import scipy.io as sio             #open matlab files module

## get data
ifile = 'data/tolpxyl.mat'       ## the path and name of the datafile

mat_contents = sio.loadmat(ifile)     ## read the file output into a dictionary

print('The contents of the mat file is ',mat_contents.keys()) ##print the names of the contents of the matfile

Aorg = mat_contents['Aorg']                        ## retrieve the Spectra from the dictionary
xorg = mat_contents['xorg']                        ## retrieve the x values from the dictionary

print('the shape of the Aorg is ',Aorg.shape)     ## inspect the variable's size


x = xorg.T                                    ## xorg is stored funny in the mat file so we transpose it
A = Aorg                               

print('x and A shapes are ',x.shape,A.shape)    
## process data



The contents of the mat file is  dict_keys(['__header__', '__version__', '__globals__', 'Aorg', 'Ctrue', 'comp', 'xorg'])
the shape of the Aorg is  (9, 601)
x and A shapes are  (601, 1) (9, 601)

Reading data out of Excel files

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

import scipy.io as sio             #open matlab files module

## get data
ifile = 'data/glass_elemental_conc.xlsx'       ## the path and name of the datafile
df = pd.read_excel(ifile,index_col=0)


print('df \n',df)    
## process data



df 
	   RI     Na    Mg    Al     Si     K    Ca    Ba   Fe  type
SID                                                                
1    1.52101  13.64  4.49  1.10  71.78  0.06  8.75  0.00  0.0     1
2    1.51761  13.89  3.60  1.36  72.73  0.48  7.83  0.00  0.0     1
3    1.51618  13.53  3.55  1.54  72.99  0.39  7.78  0.00  0.0     1
4    1.51766  13.21  3.69  1.29  72.61  0.57  8.22  0.00  0.0     1
5    1.51742  13.27  3.62  1.24  73.08  0.55  8.07  0.00  0.0     1
..       ...    ...   ...   ...    ...   ...   ...   ...  ...   ...
210  1.51623  14.14  0.00  2.88  72.61  0.08  9.18  1.06  0.0     7
211  1.51685  14.92  0.00  1.99  73.06  0.00  8.40  1.59  0.0     7
212  1.52065  14.36  0.00  2.02  73.42  0.00  8.44  1.64  0.0     7
213  1.51651  14.38  0.00  1.94  73.61  0.00  8.48  1.57  0.0     7
214  1.51711  14.23  0.00  2.08  73.36  0.00  8.62  1.67  0.0     7

[214 rows x 10 columns]

Loading Multiple Spectra and their Graphs

This idea means that you have to load more than one spectrum or chromatogram into memory for plotting. In this section we will first learn several techniques of loading multiple spectra, and then later we will plot them in several techniques. There are 3 basic strategies that can be employed to load and graph multiple spectra.

  1. manually if you just have a few spectra
  2. with a loop if you have lots of spectra or
  3. just load a *.mat file that I or someone else prepared for you.

The strategy that is used is somewhat subjective; meaning that we get to make the choice. I usually choose based upon my level of laziness at the moment.

For the demonstration of the different data loading strategies we will employee the spectral overlay graphing technique without any discussion of it. We will discuss the graphing a little later.

Multiple spectra: spectral overlays with just a few files loaded manually

There are two ways to do this. One is in x, y pairs.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
## one datafile
ifileA = 'data/sample1.csv'
rawdataA = np.genfromtxt(ifileA,delimiter=",")         ## read the file output into an array
print('the shape of the d is ',rawdataA.shape)
xA = rawdataA[:,0]                                    ## extract the wavenumbers
specA = rawdataA[:,1]                                 ## extract the absorbance values

## second datafile
ifileB = 'data/sample2.csv'
rawdataB = np.genfromtxt(ifileB,delimiter=",")         ## read the file output into an array
print('the shape of the d is ',rawdataB.shape)
xB = rawdataB[:,0]                                    ## extract the wavenumbers
specB = rawdataB[:,1]                                 ## extract the absorbance values


## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))  ## figsize is the aspect ratio of the figure
ax.plot(xA,specA,label='spec 1')        ## ax is the variable for the axis where you want to work
ax.plot(xB,specB,label='spec 2')        # do it again to plot the second spectrum 

ax.set_xlim((650,4000))                 ## set graph's limits
ax.invert_xaxis()                       ## vibrational spectra need this
ax.legend()                             ## call the legond for the axis to show it

ax.set_xlabel('Wavenumber /cm$^{-1}$')  ## latex works here
ax.set_ylabel('Absorbance')             ## latex works here
ofile = 'figures/dataspecoverlay2.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing


### This part is for making this website, so you don't need it
return ofile

dataspecoverlay2.png

Multiple spectra: loaded with a for loop (Just load and plot the data)

Notice in the code below that the data is not stored, but simply plotted. This might be useful, but does not allow the user to do anything else with all of these spectra.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

## define some custom functions
def dataloader(ifile):
    rawdata = np.genfromtxt(ifile,delimiter=",")         ## read the file output into an array

    x = rawdata[:,0]                                    ## extract/slice the wavenumbers (x values)
    spec = rawdata[:,1]   
    return x,spec

## get data file list
filelist = ['specfile1.csv', 'specfile2.csv', 'specfile3.csv', 'specfile4.csv', 'specfile5.csv']  ## list of files
datahole = 'data/'



## load and plot data
fig, ax = plt.subplots(figsize=(12,9))  ## figsize is the aspect ratio of the figure

for ifile in filelist:                        ## loop through filelist
    name = ifile[:-4]                        ## create a legion name from file list
    x,spec = dataloader(datahole+ifile)    ## use you new function to load the data;notice the concatantion of filename and path
    ax.plot(x,spec,label=name)         ## plot each spectrum; add the label 



ax.set_xlim((650,4000))                 ## set graph's limits
ax.invert_xaxis()                       ## vibrational spectra need this
ax.legend()                             ## call the legond for the axis to show it

ax.set_xlabel('Wavenumber /cm$^{-1}$')  ## latex works here
ax.set_ylabel('Absorbance')             ## latex works here
ofile = 'figures/dataspecoverlayforloop.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing


### This part is for making this website, so you don't need it
return ofile

dataspecoverlayforloop.png

Multiple spectra: loaded with a for loop (Load, store and plot the data)

There are two ways to load the data, store it in an array-like object and plot. One uses an array, and the other uses either a dictionary or dataframe.

for loop: array data storage

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

## define some custom functions
def dataloader(ifile):
    rawdata = np.genfromtxt(ifile,delimiter=",")         ## read the file output into an array

    x = rawdata[:,0]                                    ## extract/slice the wavenumbers (x values)
    spec = rawdata[:,1]   
    return x,spec

## get data file list
filelist = ['specfile1.csv', 'specfile2.csv', 'specfile3.csv', 'specfile4.csv', 'specfile5.csv']  ## list of files
datahole = 'data/'



## load and plot data
fig, ax = plt.subplots(figsize=(12,9))  ## figsize is the aspect ratio of the figure

A = np.zeros((5,3351))                 ## pre populate the array with zeros

for n,ifile in enumerate(filelist):                        ## loop through filelist; enumerate to create a counter
    name = ifile[:-4]                        ## create a legion name from file list
    x,spec = dataloader(datahole+ifile)    ## use you new function to load the data;notice the concatantion of filename and path
    ax.plot(x,spec,label=name)         ## plot each spectrum; add the label 
    A[n,:] = spec                ## store each spectrum in an array as they are loaded in the for loop


ax.set_xlim((650,4000))                 ## set graph's limits
ax.invert_xaxis()                       ## vibrational spectra need this
ax.legend()                             ## call the legond for the axis to show it

ax.set_xlabel('Wavenumber /cm$^{-1}$')  ## latex works here
ax.set_ylabel('Absorbance')             ## latex works here
ofile = 'figures/dataspecoverlayforlooparray.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing


### This part is for making this website, so you don't need it
return ofile

dataspecoverlayforlooparray.png

for loop: dictionary / dataframe data storage load file

In this technique the loading data and plotting are separated (2 for loops). They could be done together (1 for loop). This (2 for loop method) was chosen because it is easier to understand what is happening.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

## define some custom functions
def dataloader(ifile):
    rawdata = np.genfromtxt(ifile,delimiter=",")         ## read the file output into an array

    x = rawdata[:,0]                                    ## extract/slice the wavenumbers (x values)
    spec = rawdata[:,1]   
    return x,spec

## multiple files from an excel file
datahole = 'data/'
filelist = ['specfile1.csv', 'specfile2.csv', 'specfile3.csv', 'specfile4.csv', 'specfile5.csv']

## load the data
datadict = {}            ## create an empty dictionary to store the data


for ifile in filelist:                    ## for loop through the filelist
    x,spec = dataloader(datahole+ifile)
    key = ifile[:-4]                       ## get dictionary key from ifile
    datadict[key] = spec                   ## store spectrum in dictiorary 
df = pd.DataFrame.from_dict(datadict,orient='index',columns=x)   ## convert dictionary to dataframe


fig, ax = plt.subplots(figsize=(12,9))    ## figsize is the aspect ratio of the figure
#for key in df.index:                      ## loop through dataframe one index at a time  
#    x = df.columns                        ## extract the xvalues from the columns of the dataframe 
#    sp = df[df.index == key].values.T    ## get the spectrum that correponds to the indexed spectrum (key)
#    ax.plot(x,sp,label=key)              ##plot that spectrum
df.T.plot(ax=ax)                 ## plot the dataframe with index as labels on the axis ax




ax.set_xlim((650,4000))                 ## set graph's limits
ax.invert_xaxis()                       ## vibrational spectra need this
ax.legend()                             ## call the legond for the axis to show it

ax.set_xlabel('Wavenumber /cm$^{-1}$')  ## latex works here
ax.set_ylabel('Absorbance')             ## latex works here
ofile = 'figures/dataspecoverlayforloopdict.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing


### This part is for making this website, so you don't need it
return ofile

dataspecoverlayforloopdict.png

for loop: dictionary / dataframe data storage load file information from Excel

In this technique the loading data and plotting are separated (2 for loops). They could be done together (1 for loop). This (2 for loop method) was chosen because it is easier to understand what is happening. We will also get the file information from an MS Excel file. This is very sophisticated and we will not rely on this for our homework.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

## define some custom functions
def dataloader(ifile):
    rawdata = np.genfromtxt(ifile,delimiter=",")         ## read the file output into an array

    x = rawdata[:,0]                                    ## extract/slice the wavenumbers (x values)
    spec = rawdata[:,1]   
    return x,spec

## multiple files from an excel file
datahole = 'data/'
mdf = pd.read_excel('data/metadatafile.xlsx',index_col=0)  ## read file info from Excel meta data file
#mdf

## load the data
datadict = {}            ## create an empty dictionary to store the data
for ifile in mdf.index:       ## for loop through the mdf.index
    x,spec = dataloader(datahole+ifile)   ## load data
    key = ifile[:-4]    ## extract key and legend name from the mdf
    datadict[ifile] = spec            ## store the spectrum in the dictionary under the key

# the indices of the dataframe are the keys in the dicitonary (or part of the filenames)
# the columns of the df are the wavenumbers of the spectra
## all spectra must be the same size for this to work
df = pd.DataFrame.from_dict(datadict,orient='index',columns=x)     ## convert the dictionary to a dataframe



# plot the data
fig, ax = plt.subplots(figsize=(12,9))    ## figsize is the aspect ratio of the figure
#for ifile in df.index:                      ## for loop through each index
#    x = df.columns                        ## if you already have the x-values in memory then you don't need to do this step  
#    sp = df[df.index == ifile].values.T     # get the spectra out one index at a time
#    name = mdf[mdf.index == ifile].values[0][0]   ## get the legend name from mdf
#    ax.plot(x,sp,label=name)              ## plot each spectrum labeled with its name

leglist = mdf.key.values.tolist()            ## extract  the legend labels from the mdf
fig,ax = plt.subplots(figsize= (12,4))     ## create an axis and figure
df.T.plot(ax=ax)                    ## plot the dataframe with index as labels
ax.legend(leglist)                  ## call the legond for the axis to show it


ax.set_xlim((650,4000))                 ## set graph's limits
ax.invert_xaxis()                       ## vibrational spectra need this
#ax.legend()                             ## call the legond for the axis to show it

ax.set_xlabel('Wavenumber /cm$^{-1}$')  ## latex works here
ax.set_ylabel('Absorbance')             ## latex works here
ofile = 'figures/dataspecoverlayforloopdict_fromexcel.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing


### This part is for making this website, so you don't need it
return ofile

dataspecoverlayforloopdict_fromexcel.png

Multiple Spectra: from a pre-made dataset file (HDF or mat)

When you have an inconvenient number of spectra to load manually, a good choice is to load a dataset. Two good choices of dataset file formats are hdf5 files and matlab files (which are actually just hdf5 files).

In this case, these datasets need to be made either by the computer software attached to the instrument that acquired the data or separately in another program. In the latter case, this process is often a custom job, so we will not show it here, but involves looping through every spectrum, storing them in an array and then saving arrays as a file. All of the spectra must be the same size for this to work.

Often these type of data don't get a legend, because the number of spectra would make the legend size ridiculous. This number of spectra to plot is a pretty good delimiter for these two types of data.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module
import scipy.io as sio             #module for opening matlab files

## get data
ifile = 'data/tolpxyl.mat'

mat_contents = sio.loadmat(ifile)  ## read the file output into a dictionary
print(mat_contents.keys())         ##print the names of the contents of the matfile

Aorg = mat_contents['Aorg']        ## retrieve the Spectra
xorg = mat_contents['xorg']        ## retrieve the x values
print('the shape of the Aorg is ',Aorg.shape)
x = xorg.T                                    
A = Aorg                               

## process data


## plot data
fig, ax = plt.subplots(figsize=(9,6)) ## figsize is the aspect ratio of the figure
ax.plot(x,A.T)                         ## ax is the variable for the axis where you want to work


ax.set_xlim((200,340))                 ## set graph's limits
#ax.invert_xaxis()                     ## vibrational spectra need this

ax.set_xlabel('Wavelength /nm')        ## latex works here
ax.set_ylabel('Absorbance')            ## latex works here
ofile = 'figures/multidataspecoverlay.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

multidataspecoverlay.png

  • ASS10 Assignment 10 loading some data (no plotting of data required)

    NOTE: for this assignment all of the files will be in the ass10 folder not in the ass10/data folder to prevent conflict between Microsoft Windows and MacOS

    • Part 1

      Create a function that will

      • load a single data file (xy columns) from a path + filename input
      • return the x values and y values as 2 separate 1D arrays
      • Assumed that the data files will be txt files of 2 columns of data each separated by comma (a CSV format)
      • HINT: we did this in the videos
    • Part 2

      Create a function that will use the function in ass10 Part 1 to

      • load a set of data files (use the function from part 1 here) from a list of file names as the input
      • store the x-values into a 1D np.array
      • store the y-values into an 2D np.array
        • the number of rows should be the number of files that you load
        • the number of columns should be the number of y-values per file
      • return the x-values as a 1D array and the y-values a 2D array

Spectra plotting (overlayed)

This method works well if you have spectra that are either the same size as each other or different sizes (number of wavelengths). One limitation is that you overwrite the spectral data as you step through the loop. So as this code stands you can not process the data.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
def dataloader(ifile):
    import numpy as np
    rawdata = np.genfromtxt(ifile, delimiter=',')  ## read the file output into an array
    x = rawdata[:,0]                                 ## extract the wavenumbers
    spec = rawdata[:,1]                             ## extract the absorbance values
    return x, spec                              ## return the wavenumbers and absorbance values

ifile1 = 'sample2.csv'                  ## datafile named stored in variable
ifile2 = 'specfile1.csv'                ## datafile named stored in variable
datahole = 'data/'                         #location of datafiles (path)
x,spec1 = dataloader(datahole + ifile1)    ## load data file notice the concatenation of the path and datfile 
x,spec2 = dataloader(datahole + ifile2)



fig, ax = plt.subplots(figsize=(9,6))    ## create a figure and axis to plot onto
ax.plot(x,spec1,label='spec 1')            ## plot one file with label
ax.plot(x,spec2 ,label='spec 2')    ## plot another file with label  

ax.set_xlim((700,2000))                      ## change the view of the graph
ax.invert_xaxis()                             ##important for IR
ax.legend()                         ## show the legend
ax.set_ylabel('Absorbance')               ## set the y axis label
ax.set_xlabel('Wavenumber /cm$^{-1}$')    ## set the x axis label
ax.grid()                                ## turn on a basic grid to help see which bands lineup



# put a vertical line on the ax axis
# linestyles [‘solid’ | ‘dashed’, ‘dashdot’, ‘dotted’ | (offset, on-off-dash-seq) | '-' | '--' | '-.' | ':' | 'None' | ' ' | '']
# alpha is transparency 0 - 1 
ax.axvline(x=837,linestyle='--',color='g',alpha=0.5)

ofile = 'figures/dataspecoverlay.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

dataspecoverlay.png

spectra plotting (stacked)

Sometimes overlaying spectra can confuse the viewer of the data. So another way to display the data is to stack them.

Whenever you stack spectra, you need to keep somethings in mind. One, the y axis is no longer accurate, so you need to indicate this to the viewer. A customary way to do this is to add a vertical bar on the graph that indicates a specific δ y distance.

Since the data are not overlayed it is difficult for the viewer to determine if spectral features overlap. So to indicate the x position of a particular band or peak, you will need to annotate the graph. There are several ways to do this:

  • add grid lines (this works well if you are printing the graph and there are lots of bands that you want the viewer to compare)
  • add a band label with position and/or height
  • add a vertical line centered on just a few of the bands

Look in the tricks below to find the details of all of these.

Here is a basic example of a stacked plot. A constant is added to one of the y vector to move it up.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
def dataloader(ifile):
    import numpy as np
    rawdata = np.genfromtxt(ifile, delimiter=',')  ## read the file output into an array
    x = rawdata[:,0]                                 ## extract the wavenumbers
    spec = rawdata[:,1]                             ## extract the absorbance values
    return x, spec                              ## return the wavenumbers and absorbance values

ifile1 = 'sample2.csv'                  ## datafile named stored in variable
ifile2 = 'specfile1.csv'                ## datafile named stored in variable
datahole = 'data/'                         #location of datafiles (path)
x,spec1 = dataloader(datahole + ifile1)    ## load data file notice the concatenation of the path and datfile 
x,spec2 = dataloader(datahole + ifile2)


offset = 0.2
fig, ax = plt.subplots(figsize=(9,6))    ## create a figure and axis to plot onto
ax.plot(x,spec1,label='spec 1')            ## plot one file with label
ax.plot(x,spec2 + offset ,label='spec 2')    ## plot another file with label  notice the offset

ax.set_xlim((700,2000))                      ## change the view of the graph
ax.invert_xaxis()                             ##important for IR
ax.legend()                         ## show the legend
ax.set_ylabel('Absorbance')               ## set the y axis label
ax.set_xlabel('Wavenumber /cm$^{-1}$')    ## set the x axis label
ax.grid()                                ## turn on a basic grid to help see which bands lineup

ax.set_yticklabels('')   # absorbance is no longer correct so turn it off

# add a bar that indicates the absorbance scale for the two spectra
ax.annotate('0.1 absorbance',(1800,0.20),(1800,0.3),ha='center',arrowprops={'arrowstyle':'|-|'})



# put a vertical line on the ax axis
# linestyles [‘solid’ | ‘dashed’, ‘dashdot’, ‘dotted’ | (offset, on-off-dash-seq) | '-' | '--' | '-.' | ':' | 'None' | ' ' | '']
# alpha is transparency 0 - 1 
ax.axvline(x=837,linestyle='--',color='g',alpha=0.5)

ofile = 'figures/dataspecstack.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

dataspecstack.png

scatter plot of linear data

Whenever you have data with limited length (number of points < 50), it is probably better to graph it as a scatter plot than a line plot. If you have only one data set this is pretty easy.

Only one data set

Just hard code into the program the data.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
C = [1, 2, 3, 4, 5]          #  these will be the x values in the graph they are simulating concentration in ppm
A = [0.1, 0.2, 0.3, 0.4, 0.5]    #this is the simulated Aborbance values
## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4))     ## figsize is the aspect ratio of the figure
ax.scatter(C,A)                            ## ax is the variable for the axis where you want to work

ax.set_xlabel('Concentration /ppm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplescatter.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing

### This part is for making this website, so you don't need it
return ofile

examplescatter.png

Multiple sets of data

If the number of sets of data are relatively small, the you can also hard code them into the code.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data toluene and xylone
C1 = [1, 2, 3, 4, 5]          #  these will be the x values in the graph they are simulating concentration in ppm
A1 = [0.1, 0.2, 0.3, 0.4, 0.5]    #this is the simulated Aborbance values
C2 = [0.98, 2.10, 3.05, 4.09, 5.01]    # toluene
A2 = [0.11, 0.25, 0.39, 0.51, 0.691]    # toluene
## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4))              ## figsize is the aspect ratio of the figure
ax.scatter(C1,A1,label='p-Xylene')                  ## ax is the variable for the axis where you want to work
ax.scatter(C2,A2,label='toluene')

ax.legend()                                          ## turn the legend on

ax.set_xlabel('Concentration of analyte in water /ppm')               ## latex works here
ax.set_ylabel('Absorbance at 252 nm')
ofile = 'figures/examplescattermulti.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing

### This part is for making this website, so you don't need it
return ofile

examplescattermulti.png

histogram of data

Here are some final exam grades from a class of mine. Notice that the grades are also hard coded. This is a nice way to do it, because the data and the code travel together, which makes the execution of the code very reproducible.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data 
# here are some final grades of mine
g = [59.87, 65.35, 53.56, 35.73, 46.15, 59.05, 43.27, 67.27, 50.27, 46.70, 43.41, 62.61, 50.27, 66.73, 61.79, 44.78, 43.41, 50.82, 48.07, 49.17, 49.72, 57.13, 44.51, 47.53, 49.17, 70.02, 61.24, 48.35, 18.73, 62.61, 40.94, 56.85, 43.14, 55.75, 62.61, 55.48, 46.15, 43.41, 54.11]


## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4))              ## figsize is the aspect ratio of the figure
ax.hist(g)


ax.set_xlabel('Grades')               ## latex works here
ax.set_ylabel('Count')
ofile = 'figures/examplehist.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing

### This part is for making this website, so you don't need it
return ofile

examplehist.png

Tricks that can be applied to most or all of the Former Graphs

Computer Screen Inspection of Data

This is what you get as the default.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
x  = np.arange(200,400)     # these will be the x values in the graph they are simulating a UV spectrum
xc = 250                         # this is the center of the band
sigma = 10                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec)                                     ## ax is the variable for the axis where you want to work

ax.set_xlabel('Wavelength /nm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespec.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespec.png

Publication Quality Printout of Data

This is basically controlled by the output method that you choose. In the fig.savefig command the keyword dpi stands for dots per inch. The bigger the dpi number the higher resolution. If you pick a raster graphic file (tiff,jpg,png) the bigger the number after "dpi=" the more pixels per inch in the final file. If you pick a vector graphics type (eps,svg, ps) then the resolution parameter is useless. Publication quality figures usually have a dpi of 600 to 1200. The publisher will tell you what they want.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
x  = np.arange(200,400)                    # these will be the x values in the graph they are simulating a UV spectrum
xc = 250                         # this is the center of the band
sigma = 10                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec)                                     ## ax is the variable for the axis where you want to work

ax.set_xlabel('Wavelength /nm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespecpub.png'
fig.savefig(ofile,dpi=300)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespecpub.png

Thumbnail or Lab Notebook Printout of Data

Since this figure does not need to be as high quality, you can simply change the resolution parameter dpi from 300 to 75.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
x  = np.arange(200,400)                    # these will be the x values in the graph they are simulating a UV spectrum
xc = 250                         # this is the center of the band
sigma = 10                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec)                                     ## ax is the variable for the axis where you want to work

ax.set_xlabel('Wavelength /nm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespecthumb.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespecthumb.png

Beamer Presentation of Data

When you are projecting data for a conference or seminar the projectors are the limiting factor in data display. Since this figure does not need to be as high quality, you can simply change the resolution parameter dpi from 300 to 150-100.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
x  = np.arange(200,400)                    # these will be the x values in the graph they are simulating a UV spectrum
xc = 250                         # this is the center of the band
sigma = 10                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec)                                     ## ax is the variable for the axis where you want to work

ax.set_xlabel('Wavelength /nm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespecbeamer.png'
fig.savefig(ofile,dpi=150)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespecbeamer.png

Reversing the x-axis for Vibrational Spectroscopic Data

According the IUPAC standards (PAC, 2007 pg 354) for spectra, higher energy should be plotted on the left side of a spectra. For vibrational spectra with wavenumber axis tick labels running between 400 and 4000 cm-1 (low to high), traditional graphics programs plot the spectra backwards. Python provides a simple means of reversing the x-axis direction (or y-axis if need be). Here is an example:

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
x  = np.arange(1600,1800)    # these will be the x values in the graph they are simulating a UV spectrum
xc = 1656                         # this is the center of the band
sigma = 50                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec)                                     ## ax is the variable for the axis where you want to work

ax.invert_xaxis()                                   ## vibrational spectra need this

ax.set_xlabel('Wavenumber /cm$^{-1}$')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespecreversex.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespecreversex.png

Change the Line Width of the Lines so the Show Up Better

On the computer screen using python, it is relatively easy to see the lines of a graph. When you display them in a beamer presention or even in publication quality graphics, it is not always so clear. You can control the linewidth.

Here is an example controlling the linewidth in the plot command:

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

lw = 5                            # use a variable to control graphs at once

## get data
x  = np.arange(1600,1800)           # these will be the x values in the graph they are simulating a UV spectrum
xc = 1656                         # this is the center of the band
sigma = 50                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec, linewidth=lw)                                     ## ax is the variable for the axis where you want to work

ax.invert_xaxis()                                   ## vibrational spectra need this

ax.set_xlabel('Wavenumber /cm$^{-1}$')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespeclinewidth.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespeclinewidth.png

Here is an example controlling the linewidth for the whole Python or Jupyter session:

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

###these next lines write the parameters lw,lsz to plotting parameters
import matplotlib as mpl           ## this loads the moduule matplotlib as mpl
lw = 5           #this is the linewidth
lsz = 12         #this is the font size 
mpl.rcParams['xtick.labelsize'] = lsz
mpl.rcParams['ytick.labelsize'] = lsz
mpl.rcParams['font.size'] = lsz
mpl.rcParams['lines.linewidth'] = lw

## get data
x  = np.arange(1600,1800)           # these will be the x values in the graph they are simulating a UV spectrum
xc = 1656                         # this is the center of the band
sigma = 50                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(5,4))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec)                                     ## ax is the variable for the axis where you want to work

ax.invert_xaxis()                                   ## vibrational spectra need this

ax.set_xlabel('Wavenumber /cm$^{-1}$')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespeclinewidth2.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespeclinewidth2.png

Change or Control the Color of the Line

You can control the color of a line in a similar method as you control the linewidth.

Table 2: Color Correspondance table
RGBValue Short Name Long Name
(1, 1, 0) y yellow
(1, 0, 1) m magenta
(0, 1, 1) c cyan
(1, 0, 0) r red
(0, 1, 0) g green
(0, 0, 1) b blue
(1, 1, 1) w white
(0, 0, 0) k black

Here is a basic or simple example of controlling the color of a line:

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
x  = np.arange(200,400)                    # these will be the x values in the graph they are simulating a UV spectrum
xc = 350                         # this is the center of the band
sigma = 50                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(6,5))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec,'r')                                     ## ax is the variable for the axis where you want to work
ax.plot(x,spec+0.5,'k')

ax.set_xlabel('Wavelength /nm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespeccolor1.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespeccolor1.png

For finer color control, use a color tuple (3 number sequence surrounded by ( ) to get definition of the color.

Usage: (R, G, B) the numbers are between 0 and 1. All ones (1,1,1) is white, all zeros is black (0,0,0), and (1,0,0) is red.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

# use a color tuple
# this is a dictionary of colors usage cl[0] yields blue (0,0,1)
# you don't have to use this dictionary, I just like it.
# you could also redefine the keys to something more memorable
cl = {0:(0, 0, 1),#;  %blue
      4:(0, 1, 0),#;  %  bright green
      6:(0, 1, 1),#;  % cyan
      3:(0, 0, 0),#;  %black
      1:(1, 0, 0),#;  % red
      7:(1, 1, 0),#  % yellow
      2:(1, 0, 1),#;  % magenta
      5:(1, 0.5, 0.5),#;  % pink brown
      8:(0.75, 1, 0.75),#  % dull light green
      9:(0.5, 0.5, 1),#  % greyblue
      10:(1, 0.5, 0),#;  % orange
      11:(0, 0.8, 0.5),#  % bluegreen
      12:(0.5, 0, 0.6),#  % purple  
      13:(1,0.75,0),## orange
      14:(0.5,0,1),  ##also purple, dark
      15:(0.78,1,0), ##ugly green
      16:(0,0.5,1),   #blue ski
      17:(1,0,0.25),  #pink red
      18:(0.39,0.39,0.49),  #grey blue
      19:(0.02,0.42,0.1),   #forest green
      20:(0.75,0.26,0.36)  #dark red 
}


## get data
x  = np.arange(200,400)                    # these will be the x values in the graph they are simulating a UV spectrum
xc = 350                         # this is the center of the band
sigma = 50                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
color1 = cl[0]   ## get the 0th color tuple 
color2 = cl[1]   ## get the 1st color tuple

print(color1, color2)
fig, ax = plt.subplots(figsize=(6,4))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec,color=color1)                                     ## ax is the variable for the axis where you want to work
ax.plot(x,spec+0.5,color=color2)

ax.set_xlabel('Wavenumber /cm$^{-1}$')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplespeccolor1.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplespeccolor1.png

In the past there was a lot of expense associated with using color in journal/report figures, so many people avoided using color. Today these ideals are much less important. However, color should not be used just for the sake of using color. Your color choices should enhance understanding of the graphics, not just prettyfy them.

Colors and markers are attributes that can be controlled in figures. Here is a table of markers.

Table 3: Marker specification look up table
specifier Marker Type
'+' Plus sign
'o' Circle
'*' Asterisk
'.' Point
'x' Cross
'square' or 's' Square
'diamond' or 'd' Diamond
'^' Upward-pointing triangle
'v' Downward-pointing triangle
'>' Right-pointing triangle
'<' Left-pointing triangle
'pentagram' or 'p' Five-pointed star (pentagram)
'hexagram' or 'h' Six-pointed star (hexagram)

Here is an example of controlling both the color and the marker in a scatter plot.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
C = [1, 2, 3, 4, 5]          #  these will be the x values in the graph they are simulating concentration in ppm
A = np.array([0.1, 0.2, 0.3, 0.4, 0.5])    #this is the simulated Aborbance values
## process data


## plot data
fig, ax = plt.subplots(figsize=(6,5))              ## figsize is the aspect ratio of the figure
ax.scatter(C,A,marker='o',color='r',label='data 1')   ## ax is the variable for the axis where you want to work
ax.scatter(C, A+0.5,marker='+',color='r',label='data 2') ## ax is the variable for the axis where you want to work
ax.scatter(C, A+0.3,marker='x',color='g',label='data 3')           

ax.legend()
ax.set_xlabel('Concentration /ppm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplescattermakers.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resolution thing

### This part is for making this website, so you don't need it
return ofile

examplescattermakers.png

Adding Axis Labels

Use the axis methods .setxlabel('your words here') and .setylabel('your label here') to add/change the labels on axes.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
C = [1, 2, 3, 4, 5]          #  these will be the x values in the graph they are simulating concentration in ppm
A = np.array([0.1, 0.2, 0.3, 0.4, 0.5])    #this is the simulated Aborbance values
## process data


## plot data
fig, ax = plt.subplots(figsize=(6,4))              ## figsize is the aspect ratio of the figure
ax.scatter(C,A,marker='o',color='r')                 ## ax is the variable for the axis where you want to work



ax.set_xlabel('Concentration /M')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/exampleaxislabels.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing

### This part is for making this website, so you don't need it
return ofile

exampleaxislabels.png

control the font size

To control the font:

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

###these next lines write the parameters lw,lsz to plotting parameters for the whole jupyter/python session
import matplotlib as mpl           ## this loads the module matplotlib as mpl
lw = 5           #this is the linewidth
lsz = 50         #this is the font size, 50 is probably WAAAAAYYYYY too big.
mpl.rcParams['xtick.labelsize'] = lsz
mpl.rcParams['ytick.labelsize'] = lsz
mpl.rcParams['font.size'] = lsz
mpl.rcParams['lines.linewidth'] = lw

## get data
x  = np.arange(1600,1800)           # these will be the x values in the graph they are simulating a UV spectrum
xc = 1656                         # this is the center of the band
sigma = 50                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))              ## figsize is the aspect ratio of the figure
ax.plot(x,spec)                                     ## ax is the variable for the axis where you want to work

ax.invert_xaxis()                                   ## vibrational spectra need this

ax.set_xlabel('Wavenumber /cm$^{-1}$')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplesfont1.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplesfont1.png

insert a chemical structure into the graphic

  • State "TODO" from [2018-02-14 Wed 20:25]

Not done yet.

Multiple Graphs on one Figure

To add multiple graphs on one figures use the subplot command to get a different axes for each graph.
subplot(numberrows,numbercolumns,plotindex)

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
x  = np.arange(200,400)           # these will be the x values in the graph they are simulating a UV spectrum
xc = 250                         # this is the center of the band
sigma = 30                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig = plt.figure(figsize=(12,8))
ax1 = fig.add_subplot(211)      ## axis 1 for subplot 1
ax1.plot(x,spec)
ax1.set_xlabel('Wavelength /nm')               ## latex works here
ax1.set_ylabel('Absorbance')

ax2 = fig.add_subplot(212)    ## axis 2 for subplot 2
ax2.plot(x,np.sin(x))
ax2.set_xlabel('Time /min')
ax2.set_ylabel('Aplitude /V')

ofile = 'figures/examplesuplots.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplesuplots.png

Resize the Displayed Region (zoom)

Sometimes the information that you want to display is not easy to see if you show all of your data at once. In these situations it is better zoom in on the regions of interest. If you want to show where the band max is for example.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module

from matplotlib.patches import Rectangle ## for drawing the rectangle

## get data
x  = np.arange(200,400,0.1)           # these will be the x values in the graph they are simulating a UV spectrum
xc = 250                         # this is the center of the band
sigma = 30                       # %sigma*2*sqrt(2*ln(2)) is the width at half height  of the Y max of the curve
spec =  np.exp(-1*((x-xc)**2)/(2*sigma**2))   ##this will make a Gaussian curve

## process data


## plot data
fig = plt.figure(figsize=(12,8))
ax0 = fig.add_subplot(211)             ## whole specturm
ax0.plot(x,np.sin(x))
ax0.set_xlabel('Wavelength /nm')               ## latex works here
ax0.set_ylabel('Absorbance')
ax0.add_patch(Rectangle((225,0.4),width=50, height=0.6,alpha=0.5))

ax1 = fig.add_subplot(212)
ax1.plot(x,np.sin(x))
ax1.set_xlabel('Wavelength /nm')               ## latex works here
ax1.set_ylabel('Absorbance')

ax1.set_xlim(225,275)          ### control the view of lower graph
ax1.set_ylim(0.4, 1)           ### control the view


ofile = 'figures/examplezoom.png'
fig.savefig(ofile,dpi=75)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplezoom.png

Adding a Vertical Line to a Graph and Text annotations

If you want to indicate where a band is in more than one spectrum, a nice way to do this is to draw a vertical line though the graph.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module
import scipy.io as sio             #open matlab files

## get data
ifile = 'data/tolpxyl.mat'

mat_contents = sio.loadmat(ifile)                ## read the file output into a dictionary
print(mat_contents.keys())                          ##print the names of the contents of the matfile

Aorg = mat_contents['Aorg']                        ## retrieve the Spectra
xorg = mat_contents['xorg']                        ## retrieve the x values
print('the shape of the Aorg is ',Aorg.shape)
x = xorg.T                                    
A = Aorg                               
C = mat_contents['Ctrue']    ## concentration matrix
## process data


## plot data
fig = plt.figure(figsize=(12,8))  ## figsize is the aspect ratio of the figure
ax1 = fig.add_subplot(211)       ## ax1 is the variable for the axis where you want to work      
ax1.plot(x,A.T)                                     

ax1.set_xlim((200,340))                             ## set graph's limits

# put a vertical line on the ax axis
# linestyles [‘solid’ | ‘dashed’, ‘dashdot’, ‘dotted’ | (offset, on-off-dash-seq) | '-' | '--' | '-.' | ':' | 'None' | ' ' | '']
# alpha is transparency 0 - 1 
ax1.axvline(x=222.5,linestyle='--',color='k',alpha=0.9)
ax1.axvline(x=274,linestyle='--',color='r',alpha=0.9)

ax1.text(224, 2.62, ' this band is not linear with concentration')  ## text annotations
ax1.text(280,1.3, '$\leftarrow$ this band  is linear with concentration')  ##latex annotations

ax1.set_xlabel('Wavelength /nm')               ## latex works here
ax1.set_ylabel('Absorbance')

ax2 = fig.add_subplot(212)
ax2.scatter(C[:,0],A[:,304],color='r',label='274 nm')
ax2.scatter(C[:,0],A[:,509],color='k',label='222 nm')
ax2.set_xlabel('Analyte Concentration v/v /%')               ## latex works here
ax2.set_ylabel('Absorbance')
ax2.legend()

ofile = 'figures/exampleverticalline.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

exampleverticalline.png

adding grid lines

Grid lines also allow your reader to interpret the data in a graph. They are particularly useful for ovelaping data sets.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module
import scipy.io as sio             #open matlab files

## get data
ifile = 'data/tolpxyl.mat'

mat_contents = sio.loadmat(ifile)                ## read the file output into a dictionary
print(mat_contents.keys())                          ##print the names of the contents of the matfile

Aorg = mat_contents['Aorg']                        ## retrieve the Spectra
xorg = mat_contents['xorg']                        ## retrieve the x values
print('the shape of the Aorg is ',Aorg.shape)
x = xorg.T                                    
A = Aorg                               

## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))              ## figsize is the aspect ratio of the figure
ax.plot(x,A.T)                                     ## ax is the variable for the axis where you want to work


ax.set_xlim((200,340))                             ## set graph's limits

# which ---> major, minor, both
# Turn on the minor TICKS, which are required for the minor GRID
ax.minorticks_on()

ax.grid(True, which='major')                ##select which to show, 'major' grid lines or 'minor'

ax.set_xlabel('Wavelength /nm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/examplegrid.png'
fig.savefig(ofile,dpi=300)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

examplegrid.png

adding a band label

Another way of indicating areas of the graph that you want to talk about in the text or bring into focus for the user is to add textual annotations such as a label to your graph. In python you can do this by importing your graph into powerpoint and manually add the label. This works fine until your professor makes you recreate every figure using the Greek alphabet. If you are finding that you need to recreate lots of figures and/or don't like using a mouse you can also add annotations programatically. Here is an example.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module
import scipy.io as sio             #open matlab files

## get data
ifile = 'data/tolpxyl.mat'

mat_contents = sio.loadmat(ifile)                ## read the file output into a dictionary
print(mat_contents.keys())                          ##print the names of the contents of the matfile

Aorg = mat_contents['Aorg']                        ## retrieve the Spectra
xorg = mat_contents['xorg']                        ## retrieve the x values
print('the shape of the Aorg is ',Aorg.shape)
x = xorg.T                                    
A = Aorg                               

## process data


## plot data
fig, ax = plt.subplots(figsize=(12,9))              ## figsize is the aspect ratio of the figure
ax.plot(x,A.T)                                     ## ax is the variable for the axis where you want to work


ax.set_xlim((200,340))                             ## set graph's limits

ax.text(275,1.2, '(a)',fontsize=25)  ## this fontsize is too big 

ax.set_xlabel('Wavelength /nm')               ## latex works here
ax.set_ylabel('Absorbance')
ofile = 'figures/exampletext.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

exampletext.png

adding a scale bar to the graph

I don't really have a great way of doing this. I will think about it and add it later.

# many functions are avaible in modules or libraries
# in this example we will load the numpy module of functions
import numpy as np                 #this command loads all of the functions in numpy and labels them np
import pandas as pd                # data organization module  
import matplotlib.pyplot as plt    # plotting module


## get data
def dataloader(ifile):
    import numpy as np
    rawdata = np.genfromtxt(ifile, delimiter=',')  ## read the file output into an array
    x = rawdata[:,0]                                 ## extract the wavenumbers
    spec = rawdata[:,1]                             ## extract the absorbance values
    return x, spec                              ## return the wavenumbers and absorbance values

ifile1 = 'sample2.csv'                  ## datafile named stored in variable

datahole = 'data/'                         #location of datafiles (path)
x,spec1 = dataloader(datahole + ifile1)    ## load data file notice the concatenation of the path and datfile 




fig, ax = plt.subplots(figsize=(12,9))    ## create a figure and axis to plot onto
ax.plot(x,spec1,label='spec 1')            ## plot one file with label

ax.set_xlim((700,2000))                      ## change the view of the graph
ax.invert_xaxis()                             ##important for IR
ax.legend()                         ## show the legend
ax.set_ylabel('Absorbance')               ## set the y axis label
ax.set_xlabel('Wavenumber /cm$^{-1}$')    ## set the x axis label


ax.annotate('0.05 abosorbance',(1750,0.0),(1750,0.05), arrowprops={'arrowstyle':'|-|'},ha='center')

ofile = 'figures/dataspeccalebar.png'
fig.savefig(ofile,dpi=100)      ## dpi is dots per inch and is a figure resotution thing
### This part is for making this website, so you don't need it
return ofile

dataspeccalebar.png

ass11

  • create a Jupyter notebook (use the blank ass11.ipynb already in the directory) from scratch in the ass11 directory
  • load one or more data file(s) from your ass11/data directory
  • plot the data properly
  • use 2 of the following tricks
    • make a stacked plot
    • added grid lines
    • use subplots to create 2 axes
    • add a bandlabel
    • add a vertical line to highlight one band
  • full credit will only be given to graphs (located in the figures directory) that are perfect, partial credit will be given.
  • No credit without showing your work (code) in the ass11.ipynb
  • make a 150 dpi graph of the data
  • save this file into the ass11/figures directory

Other resources

Date: Spring 2023

Author: Scott Huffman

Created: 2023-01-31 Tue 10:18

Validate