CHEM455/CHEM555: Chemometrics
Table of Contents
- Class information
- First day lecture and some motivational words
- What is Chemometrics?
- Examples of Chemometrics: Concentration Determination.
- How do you deal with mixtures and Beer's law?
- Examples of Chemometrics: What is it?
- Identification of unknown material: IR spectral classification
- What is it?
- Example questions from Forensic Science
- Chemometrics is
- Industry standard options
- Why not MS Excel? So you noticed that MS Excel was not listed above as a tool for doing chemometrics?
- Octave is a free choice
- R, S, SAS, SPSS
- Python is a modern choice
- Getting the software and assignments
- Introduction to Python in Jupyter
- Introduction to Jupyter
- A computer is a calculator
- Functions and a few commands
- Some more Jupyter knowledge (Markdown)
- User input
- Control Structures: if
- Strings, Lists
- Strings
- Concatenation
- indexing and strings
- slicing stings
- Learning Review
- Learning Review: Answers
- the immutable string
- determine the length of a string
- Learning Review A:
- Learning Review B: More Advanced (to do at your own pace)
- lists
- Defining lists
- indexing and slicing lists
- Concatenation of lists
- Lists are Mutable
- Adding items to the end of a list
- Getting the length of a list
- Lists of Lists
- learning Review
- Assignment 6 (ass6) Part 1: create a function that
- Assignment 6 (ass6): Part 2 List slicing (10 points)
- control structures: for loop
- Arrays and Matrices for data storage
- Dataframes (the named arrays)
- Making Graphics and Figures
- General Philosophy for what type of graph to use
- Generic Example
- plotting useless data
- First some data
- Loading Multiple Spectra and their Graphs
- Multiple spectra: spectral overlays with just a few files loaded manually
- Multiple spectra: loaded with a for loop (Just load and plot the data)
- Multiple spectra: loaded with a for loop (Load, store and plot the data)
- for loop: array data storage
- for loop: dictionary / dataframe data storage load file
- for loop: dictionary / dataframe data storage load file information from Excel
- Multiple Spectra: from a pre-made dataset file (HDF or mat)
- Spectra plotting (overlayed)
- spectra plotting (stacked)
- scatter plot of linear data
- histogram of data
- Tricks that can be applied to most or all of the Former Graphs
- Computer Screen Inspection of Data
- Publication Quality Printout of Data
- Thumbnail or Lab Notebook Printout of Data
- Beamer Presentation of Data
- Reversing the x-axis for Vibrational Spectroscopic Data
- Change the Line Width of the Lines so the Show Up Better
- Change or Control the Color of the Line
- Adding Axis Labels
- control the font size
- insert a chemical structure into the graphic
- Multiple Graphs on one Figure
- Resize the Displayed Region (zoom)
- Adding a Vertical Line to a Graph and Text annotations
- adding grid lines
- adding a band label
- adding a scale bar to the graph
- ass11
- Other resources
Class information
First day lecture and some motivational words
What is Chemometrics?
My Definition of Chemometrics
The application of numerical methods to extract chemical information from chemical data.
What are numerical methods?
Math
What is chemical Data?
numbers that you get from measurements
What is chemical information
- results
- constants
- equilibrium constants
- rate constants
- bond lengths
- knowledge
- models
- conclusions??
- etc.
Another resource
Examples of Chemometrics: Concentration Determination.
All analytical chemistry deals with mixture analysis. From this perspective there are two issues that are coupled.
- What are the compounds in the sample?
- How much of each are there?
- This first part deals with "How much is there?"
Determining concentration of an unknown sample.
Let us pretend you are a chemist and need to know the concentration of phosphate in an aqueous solution. What do you do?
- First is the calibration standards (external calibration)
- prepare a series of phosphate solutions of known concentrations
- measure the UV absorbance spectrum of each solution.
- WAIT A MINUTE!!!!
- you don't measure absorbance, what do you measure?
- if you are using a photocurrent detector for your measurement,
- then you measure current as a function of wavelength (sometimes called a single beam spectrum) of each solution including from a blank sample
- typically the computer software of you spectrometer converts single beam spectra of each solution and the blank into an absorbance spectrum using the following equation
- Absorbance Spectra of Phosphate
Figure 1: Absorbance spectra of sodium phosphate solutions
- Create table of concentrations of phosphate and band absorbance
Table 1: Phosphate data Concentration of Phospate /mM Absorbance at 210 nm 0.100000 0.097067 0.300000 0.296739 0.500000 0.468451 0.700000 0.669580 0.900000 0.861085 Note: 210 nm is the top of the spectral band for phosphate. This is the analytical wavelength.
- Make a Beer's Law Graph
Figure 2: Beer's Law graph of phosphate absorbance at 210 nm vs. known concentration of phosphate
- Fit the data with a linear function
Because of the linear relationship between concentration of photon absorbing species and the number of photons absorbed. (This relationship is called Beer's Law.)
\begin{equation} A = \epsilon l C \end{equation}we can fit the data with a linear function.
\begin{equation} A = m C + b \end{equation}This equation (specifically the slope and intercept) is a model for the behavior of this chemical system (chemical information).
- m is the slope (sensitivity of the photon absorbing analytes to absorb light)
- b is the y-intercept (empirical parameter that encapsulates the optical differences between the blank and the calibration standards, ideally this is zero.)
- these (m and b) are also called the fitting parameters
"Measure" the aborbance spectrum of the solution of unknown concentration
- and get its absorbance at 210 nm. Remember this wavelength (210 nm) is the analytical wavelength, because we are using the absorbance at this wavelength to determine concentration.
- in this case it is also the λmax, which is the wavelength of maximum absorbance
- solve the linear function for C
- and plug the unknown solution's absorbance at 210 nm into the rearranged linear function
Cunk is the concentration of phosphate in the solution, and is also chemical information.
What assumptions were made during this example?
- unknown solution has the same chemical environment as the calibration standards
- interactions between absorbing species do not change with concentration
- usually means: relatively low concentrations (therefore we can ignore Activity)
- no interaction between absorbing species and matrix
- usually means: don't have to use method of standard addition
- the analyte (phosphate in the example above) is the only absorbing species at the analytical wavelength
- Very simple system,
- requires lots of unknown sample preparation to get the solution to only contain phosphate
- lots of human labor for highly variable situations (blood, river water, detergent factory run-off, drainage near a mining operation, etc.)
- and therefore of only limited use
- or involves chromatography to cleanup the sample and takes a "long" time per sample
In the example above, the sample and calibration standards contain phosphate PO43+ + Na+ in water. It is a mixture, but only the phosphate anion absorbs UV light. So from a spectroscopic/chemometric perspective this chemical system is a single component system.
How do you deal with mixtures and Beer's law?
Another word for mixture analysis is multicomponant analysis. There are 2 situations for calibration curve creation and unknown concentration determination.
1: You know every component's pure spectrum and concentrations
| A | B |
|---|---|
| spectra don't overlap | spectra do overlap |
| can use 2 curves | ??? |
2: You do not know every component's pure spectrum nor do you know all of the concentrations of all of the species that contribute to the spectra
In other words, you know the concentration of your analyte in the calibration standards but not the concentrations of other spectral contributors in you unknown samples, and maybe even your calibration standards.
This is a bad situation that is all too common in the real world. You might be asking yourself, why does this situation occur? Here are some reasons:
- the sample that you need to analyze is precious (think Gollum from the Hobbit and Lord of the Rings) and can not be replaced or damaged in the measurement. Examples
- forensic science evidence
- art that is not replaceable
- the sample is expensive
- purifying the sample renders the sample useless
- Here are two examples from the literature
- the sample as a mixture has important value that is greater than the purified parts (think BBQ sauce, or steak)
- at the current technological state of the world, we can not simply mix the chemical components in wine or beer, and call the mixture good wine or beer
- instead we have to grow these products
- so identifying all of the components in the beer may not allow you to predict its quality better than measuring the beer as a whole
- https://doi.org/10.1080/10408398.2012.726659
- https://doi.org/10.1016/j.aca.2006.04.070
Some examples that are real but don't fit into the categories above
- NIR spectroscopic determination of glucose concentration in blood (everything in the blood contributes to the NIR spectra of blood)
- but for a diabetic knowing the glucose concentration in their blood is critically important for maintaining their health
- UV spectroscopic determination of nitrate in natural waters (nitrate, phospate, and chloride all "look" similar to each other in the UV spectra)
Figure 6: Toluene and p-xylene UV absorbance spectrum in a series of 2 component mixtures
You will learn in this course, how to predict the concentration of multiple components in mixtures of both situations 1 and 2.
Examples of Chemometrics: What is it?
All analytical chemistry deals with mixture analysis. From this perspective there are two issues that are coupled. What are the compounds in the sample, and how much of each are there. We saw some examples of the second part of this sentence earlier. These types of tasks are sometimes called quantitative analysis. This first part of the sentence deals with "What is this unknown material, and how does it relate to other materials?" In other courses this task is sometimes called qualitative analysis.
Identification of unknown material: IR spectral classification
In organic chemistry you probably did some IR spectral analysis in which you identified parts of a compound by its IR spectral features. For example, you might have been given the following IR spectrum and asked to determine if it contains a carbonyl functional group or not, and if so what kind of carbonyl (i.e. amide, ketone, carboxylic acid, aldehyde, etc)
Figure 7: What kind of molecule is it?
What is it?
Lots of analytical tools will give us data that can be used to determine chemical identity.
Example questions from Forensic Science
- What polymer was used in this textile?
- Is this powder cocaine?
- Is this powder an explosive residue?
- Is this powder anthrax spores?
These things can be done by visual inspection of the data. An example is spectral interpretation, and manual comparison of bands/peaks in a spectrum to a set of reference tables. These both take lots of time.
- Washington, D.C. anthrax spores were sent in the mail.
- Ideally: need to scan every letter that passes through the mail system (millions per day)
- Whenever we need to repeat a task many, many times, often it is better to let a computer do it.
- The trick is figuring out how to train the computer to do what a human can do.
- This is what chemometrics is all about.
Chemometrics is
- getting the computer to process your data for you, rather than processing it by hand.
Therefore we will need a computer and software to do this.
Industry standard options
Why not MS Excel? So you noticed that MS Excel was not listed above as a tool for doing chemometrics?
- you are a scientist
- MS Excel is a tool for buisness people who are afraid of doing math
- I am a MS bigot
- spreadsheets are not good for massive amounts of data
- So this is a choice. You can make your own, but I am going to use another choice for this course.
Octave is a free choice
- similar to MATLAB (mostly compatable with Matlab)
- free
- open source
- if your really like it you can contribute to it or help maintain its awesomeness
R, S, SAS, SPSS
- relatively easy to use
- these are called area specific languages (they do stats well, but only stats)
- little difficult to understand what is happening underneath the hood
- I prefer something else, and I am the teacher
Python is a modern choice
- One limitation of all of these other choices is that they are all area specific tools
- this can be a limitation, if you ever want to do something else besides process data
- Python is a general purpose computer language
- it can be used for most computer programming tasks
- it is becoming, if not already is, the language for Data Science
- Chemometrics is a subset of Data Science (and pre-dates it by more than 40 years)
- this is what we will use for this class
Getting the software and assignments
Anaconda distribution
We will use the Anaconda distribution for getting python in this course. Some benefits are
- all self-contained
- easy to install all of the parts
- does not interfere with the rest of your computer
You have 2 choices:
- use the http://virtual.wcu.edu (requires internet/network access to use, but very little computer resources)
- install anaconda
- goto https://www.anaconda.com/download/
- download the Anaconda installer for your computer (Linux, MacOS, Windoze) Python version should be > 3.6 (Python 3.9 as of 2022)
- install it on your computer
- run it to make sure that it is working
Here is a pdf with screen captures of me installing the Anaconda distribution on a Windoze computer: ./figures/anaconda_installation_running_guide.pdf
ass01
Prove to me that you able to run python. Go to your menu system and start Jupyter notebook. This will open a tab in your default browser. In Windoze, it should look something like:
./figures/anaconda_installation_running_guide.pdf
- For now, select Python 3
- This should open a new tab in you browser, running jupyter/python
- Do the following in Jupyter Notebook
Assignment 1: 1 point
In the Cell beside In [ ]
Type the following things
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt
and press Enter on your keyboard while holding down the SHIFT key
print('Your Name')
and press Enter on your keyboard while holding down the SHIFT key
Screen capture the Juypter tab.
Upload this image file in the Learning Management System DropHole under ass01 or "Proof of Installation"
- You can delete this Jupyter notebook file after you email the image.
Dropbox hole
In this course, we will use dropbox to share files. So your second assignment is to get a Dropbox account (if you already have one you can use that one). With this account is an associated email address. It may be a catamount email address or it may be something else. Send me an email from this dropbox associated account. Put in the Subject line:
Assignment 2: Dropbox
I will then share a dropbox folder with you. Once my account and your account are linked in this directory, you should see in that directory/folder the course files that we are sharing. Note: I will not be able to see any other dropbox stuff of yours but only this one directory. We will refer to directory as the course home directory. Make sure that when you run Jupyter that you can access this course home directory. (One way to do this is to run Jupyter notebook/lab from this directory.) Another way is to put the Dropbox directory within your course directory on your computer. See the following structure.
Figure 9: Example of where to put your Dropbox on a Windoze computer
chemoclass
├── Dropbox
├── learn
├──resources
├──ass01
In this course home directory will be several subdirectories:
- resources - this is the textbook, and papers to read
- learn - this is where you can practice your stuff
- assignments - there will be lots of these that correspond to the assignments. Each assignment will have a unique directory. They will begin with:
- ass## - assignment directory the ## is 01, 02, etc that corresponds to the assignment number
An example homework directory may already be in the dropbox course home directory, called ps1. You can poke around there if you want. This directory will never be accessed by your instructor.
Here is what the course home directory will look like.
├── learn <-- practice stuff here ├── ps1 <-- example assignment └── resources <-- reading material (the text book in pdf form)
The learn directory
The learn directory will have the following structure at the start. This is were you can do most of your learning/practicing for this course.
learn ├── data ├── figures ├── lib └── notebook_template.ipynb
The whole enchilada
Here is what your dropbox course home should look like the at the beginning of this course.
├── learn
│ ├── data <-- this is where learning/practicing data that you will process is located
│ ├── figures <-- this is where you save your figures generated in python
│ ├── lib <-- in this directory are library files that we will use later
│ │ ├── analtool3.py
│ │ └── chemo.py
│ └── notebook_template.ipynb
├── ps1 <-- practice assignment
│ ├── jupyter.png
│ ├── problem1.ipynb
│ └── problem2.ipynb
└── resources
└── textbook.pdf
Assignments
Whenever you work on an assignment, make sure that you don't change the name of the files, and make sure that the files are in the same directory from which they came.
Once the assignment has beed graded, you will receive feedback in the same directory as the assignment and/or via email.
ass02
Assignment 02 will be proof that you have installed dropbox on your computer and can receive an assignment and submit it back to me. In the ass02 directory open the jupyter notebook and type your name in the cell below the instructions.
Introduction to Python in Jupyter
Here is a tutorial for python. Much of the ideas in this learning module can be found at this tutorial.
Introduction to Jupyter
Whenever you run Jupyter on your computer, you will see in your browser the following:
From this screen you can either start a new python notebook or open an existing one. In your home/learn directory you will find at least one notebook template. Copy that template to you working Jupyter directory. From the home screen you can check the notebook template checkbox and an option at the top to duplicate the notebook will appear. Do this. The reasoning is that you preserve the template for future use and only work on copies of it. You can rename the copy to anything that you like.
Figure 10: Jupyter notebook page running
To run python code, you type the code in to an active cell in Jupyter (as indicated by the blue bar on the left side of the active cell). Once you have finished typing the code to execute (or run) the code you simultaneously depress the Shift-Enter keys. (hold SHIFT and press ENTER)
Use the template
A lot of our code in this course will rely upon pieces of code that other people have written, these pieces of code are stored in modules that come with the Anaconda python distribution that you have hopefully already installed. A template has been provided in your learn directory. You can start with this template and make a copy of it directly from Jupyter. See if you can figure out how. I would suggest that you keep a pristine copy of this template. That way if you break what you are working on, you can go back.
learn ├── data ├── figures ├── lib └── notebook_template.ipynb <------HERE
A computer is a calculator
For the purposes of this course and in my opinion, a computer is nothing more than a calculator (although a big, powerful, awesome one). (parts 1-5)
What do you do with a calculator?
- read/write data
- process data
- plot data
- evaluate data
- extract information from data (math)
What are the parts of the calculator
- user interface for input (for us: jupyter)
- user interface for output (for us: jupyter)
- a computational engine for doing the math (for us: python)
- storage space for saving things for later usage (for us: jupyter for temporary and your hardrive for longer term)
comments in python
- Use the # symbol to add comments (extra info that might be useful latter but is not used in the calculations) to your math or code
- definition: code are the words and other groups of letters and numbers that you type to get the computer to do work
Some Rules of python
- the code in python is case sensitive, so a variable named tup is different than Tup, and TUp and TUP, etc.
- some symbols have specific predefined meaning in python:
- all variables must begin with either a letter or an underscore
- after the first character the remainder of your variables may consist of any number of letters, numbers or underscores
- Don't use the following words in variable names that you create
| and | assert | break | class | continue |
| def | del | elif | else | except |
| exec | finally | for | from | global |
| if | import | in | is | lambda |
| not | or | pass | raise | |
| return | try | while | ||
| False | class | finally | is | return |
| None | continue | for | lambda | try |
| True | def | from | nonlocal | while |
| and | del | global | not | with |
| as | elif | if | or | yield |
| assert | else | import | pass | |
| break | except | in | raise |
or these
| numpy | Float | Int | Numeric | Oxphys |
| array | close | float | int | input |
| open | range | type | write | zeros |
or these
| acos | asin | atan | cos | e |
| exp | fabs | floor | log | log10 |
| pi | sin | sqrt | tan |
- white space matters: don't use spaces in filenames or functions or variables
- don't use hyphens - or parentheses or brackets ( or ), [, ], {, or ] in variable names
- don't use any special character other than underscores invariables names
All of these special characters have specific meaning in python and will be mis-interpreted
- Review:
- type some words from the tables above in the cells of jupyter to see that the change color (this is called syntax highlighting), and it means that Jupyter recognized those words as predefined Python words
- type some variables in a different cell disobeying the rules (just do one at a time) and execute the cell (Shift-ENTER or Shift-Return). Look at the error messages that are created.
- Use this as a template for defining a variable (x is the variable).
x = 1.0
displaying stuff
- use print( whatever you want to print, separated by commas) to print stuff to the screen
arithmetic
Computers can do arithmetic. Here are the commands for various arithmetic operations.
Type the following things in a jupyter cell and press shift-enter after each one:
print(4 + 4) # this command prints the results of 4+4 to the screen 4+4 #this one does not (unless it is the last command in the cell) print(4 - 4) #What does this do? print(4 * 4 ) #and this print(4 / 4) print(4**2)
8 0 16 1.0 16
in these examples the +, -, , /, and * are called operators.
Note: In Python ** is the operator symbol for power.
algebra
Lets say we have this simple algebra equation
\begin{equation} 4=x-1 \end{equation}and your task is to solve for x. What do you do? You subtract -1 from both sides to give you
\begin{equation} x = 4+1 \end{equation}in this equation x is called a variable. In a cell, type:
x =5 #this command assigns the value 5 to the variable x # but will not print the contents of x print(x) #this command prints the contents of x
5
Now type in a cell
x = 4+1 # this command assigns the results of the operation 4+1 to x # but will not print the contents of x print(x) # ths command prints the value contained in x
5
Notice how you get the same results. This means that the computational engine (python interpreter) is computing 4+1 and storing it into the variable x.
Note for later: 5 is called a scaler number.
Once you have this variable filled you can perform operations on it. Type
x = 5 print(x*5) ##this command prints x times 5 to the screen
25
or we store this result in a new variable
x = 5 u = x*5 print(x,u) #this command prints x and u to the screen
5 25
Fancy printing part 1
x = 5 u = x*5 print(x,u) #this command prints x and u to the screen print('x*5 = ',x) #this command prints x to the screen with some text showing how it was calculated
5 25 x*5 = 5
Fancy printing part 2
The previous output may not necessarily be what you want. Take this example:
# pi = 3.141592653589793238462643383279502884197169399375105820 print('pi = ',pi) # this prints an undefined by you amount of digits print('pi = {:.2f}'.format(pi)) # in this case you choose the significant digits to print
pi = 3.141592653589793 pi = 3.14
This example uses the format function. It does 2 things.
- it converts the number stored in the pi variable to a string (letters and numbers and symbols) for display/printing purposes
- it allows you to control what that string looks like
Why do this? Humans and computers need different formatting to read information. This format command converts computer looking information to human useful formatting.
In the second print command the f tells the computer that the number (pi) is a f*loating point number (has a decimal point). The *.2 tells the computer that you want 2 digits to the right of the decimal.
Here is a link for how to control this formatting. You can explore this at your leisure.
- Review:
- Store you name as a string variable
- Figure out how old you are in Days (http://jalu.ch/coding/days/en)
- Store this number in a variable
- Calculate the percentage of your living days in a century of days ( you can use 36525 as an approximation) and store it as a variable.
- print a sentence (string using single quotations) that says something like: My name is X, and I have lived Y percent of a century.
- Use the format method (command) to substitute your name and century percentage in the printed sentence
geometry
Given the figure below what is the distance r?
Figure 11: Distance in 2D space
You may or may not remember the Pythagorean theorem
\begin{equation} r^2 = a^2 + b^2 \end{equation}where
\begin{equation} a = 3 -1 \end{equation}and
\begin{equation} b = 2-4 \end{equation}To solve for r, we need to take the square root of each side of the equation above.
\begin{equation} r = \sqrt{a^2 + b^2} \end{equation}Let us calculate r using python. Type
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np a = 3-1 b = 2-4 r = np.sqrt( a**2 + b**2 ) ## note the np.sqrt, note the ** for power print('np.sqrt( a**2 + b**2 ) = ',r) print('np.sqrt( a**2 + b**2)= {:2.2f}'.format(r))
np.sqrt( a**2 + b**2 ) = 2.8284271247461903 np.sqrt( a**2 + b**2)= 2.83
What is the np.sqrt word in python? It is a built-in function; meaning that it comes with python as part of the numpy module that we loaded as np. Functions have the format
output = function( input ) Note the type of parentheses
where output (called returns in the python lingo) is the result of the function doing its thing, and input is the directions, data, and numbers that the function uses. This function np.sqrt computes the square root of a number.
Probably every mathematical function that you might need is available in the numpy module.
- General ideas about functions
- they can have more than one input (called parameters in the python lingo)
- the inputs can be any kind of data structure, i.e. numbers, letters, matrices, vectors, etc.
- to assign output to a variable simply used an equal sign =
- function usage: outputvariable = function(input)
- help about a function can be found by typing a ? in front of or after the function
type ?np.sqrt in a cell and execute the cell.
Review: Look at the numpy mathematical functions reference page
https://numpy.org/doc/stable/reference/routines.math.html
- lookup some functions (mathematical) that might be useful and figure out how to use them. Here is a list that I use all the time
- cos
- sin
- tan
- round
- exp
- log
- log10
- sqrt
- max
- min
- Check you answers in python with a calculator to ensure that you are using them correctly.
- Here is an example (https://numpy.org/doc/stable/reference/generated/numpy.round_.html#numpy.round_)
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np pi = np.pi ## notice that we also get pi from the numpy (np) module xrounded = np.round(pi, 5) ### round pi to 5 decimal places print(xrounded)
3.14159
Functions and a few commands
Showing the cell line numbers
If you want to show the cell linenumbers, click to the left of the [ ]: of any cell. You should see a thick vertical blue bar. Then depress Shift-L. This key-stroke sequence will toggle the cell line numbers on and off.
Some more python/jupyter info
First some more ideas about how jupyter/python
- Each time you start jupyter is called a session. By default all of the memory is wiped for each session.
- Each time you execute a command block or cell (type the command and depress the shift-enter keys) the results are stored in the memory for that session.
type each in a cell and execute:
x = 1 who ##this command prints the variables currently in memory
y = 100 whos # this is a "magic" command and prints variables and their types
Building your own functions (ideal setup)
The numpy module for python contains probably every math function that you might want. Here is the website. https://numpy.org/, so you can check it out.
What if numpy or Python does not have a function that you want? What do you do?
Build you own!
For example, if we wanted to find the angle Q in the geometry problem from 11. You would need to use the following trigonometric relationship
\begin{equation} Q = arctan\left( \frac{a}{b} \right) \end{equation}How do you go about creating a function to find Q?
First solve the problem (for now, not a function)
Before we create a function lets programmatically solve the problem, then wrap the solution into a function.
Let us solve the angle problem that we saw before and find Q. Type the following in a Jupyter cell and execute
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np a = 3-1 b = 2-4 Q = np.arctan(a/b) print('the angle Q is equal to {:2.3f}'.format(Q))
the angle Q is equal to -0.785
OK, so this gives you a benchmark in which to compare, i.e. the correct value for Q. Our function should give us the same answer. Note that the angle is in radians not degrees. This is the default.
OK now a function
Defining a new function in python involves wrapping the code with
def function_name( inputvariable1, inputvariable2): ## end the first line of a code block with a colon ''' this is a comment section or help section that prints when you type ? before a function in jupyter notice the indentation. this is important in python ''' import numpy as np ### this is where you type the code for your calculations return outputvariable ##return stuff back to the caller
Note: In python, blocks of code are delineated by indentation
Now wrap the findq code above into a function.
def findq(): ## end the first line of a code block with a colon ''' for now let's just have the function make the calculation and print it out. Later we will make it more flexable ''' import numpy as np #this command loads all of the functions in numpy and labels them np a = 3-1 #Coordinates hard coded into the function b = 2-4 #Coordinates hard coded into the function Q = np.arctan(a/b) print('the angle Q is equal to {:.3f}'.format(Q)) return #notice the indentation is back to the first column. So the above code block is ended findq() # you need the () they are part of the function.
the angle Q is equal to -0.785
Note: when you define a function you have not run or executed or called the function. To do that you have to use the function's name or call the function. (See the last line in the code block above.)
What does this function do? It calculates the angle Q from a set of fixed coordinates that are coded into the function and prints the value of Q to the screen. This approach has several limitations:
- you need to redefine the function for a different a, b, and Q. (type in the triangle's coordinates for a and b and execute it (Shift-ENTER)
- To use the value of Q in this form you would need to copy the value by hand and type it again (or copy->paste)
- The value has been truncated. That could give you rounding errors further down the line of calculations. (Remember you don't round until the end of you calculations.)
Now a more useful function
Great. Now let us make this function a little more readable and useful. In the function findq
- add some comments to the header portion of the function
- some commented definitions to explain what this function does.
- return the value of Q to the caller
def findq(): ## end the first line of a code block with a colon ''' this function calculates the angle between to line segments. Usage: angle = findq() prints the answer to the screen and returns the angle The angle has units of radians ''' import numpy as np #this command loads all of the functions in numpy and labels them np a = 3-1 b = 2-4 Q = np.arctan(a/b) print('the angle Q is equal to {:2.3f}'.format(Q)) return Q #return angle Q angle = findq() ##store q in the variable angle print('The square of the angle is ',angle**2, 'and the angle is',angle) #now we can use this angle for something else
What does this function do?
- It calculates the angle Q from a set of fixed coordinates that are coded into the function
- prints the value of Q to the screen.
- it returns the value of Q with all of its digits (that can be utilized in future calculations)
What are its limitations:
- it only calculates the angle of a very specific triangle
and now an even more useful function
This function findq is not that useful; it only calculates the angle of a very specific triangle, prints the value to the screen, and outputs the angle.
You could do the same thing a lot easier without a calculation:
def sameasfindq(): ''' same behavior as findq ''' print('the angle Q is equal to -0.785') return -0.785398163397 angle = sameasfindq() print('The square of the angle is ',angle**2, 'and the angle is',angle)
the angle Q is equal to -0.785 The square of the angle is 0.6168502750673807 and the angle is -0.785398163397
It is not quite the behavior that we want. A better function would be one that could be used for multiple problems i.e. a more flexible function.
Let's alter the function to make it one that can calculate the angle given by a user's set of line segments. NOTE: Who is this user? It could be you. Or a future you. Or a collaborator to whom you share the function?
Users of the function tell function what to use for input by specifying the variable inside the () in the definition.
def findq(a,b): ## end the first line of a code block with a colon ''' this function calculates the angle between to sides of a triangle. Usage: Q = findq(a,b) prints the answer to the screen and returns the angle The angle given has units of radians a is length of the opposite side b is the length of the adjacent side ''' import numpy as np #this command loads all of the functions in numpy and labels them np #a = 3-1 #commmented out so you can see the change #b = 2-4 #commendted out so you can see the change Q = np.arctan(a/b) print('the angle Q is equal to {:2.3f}'.format(Q)) return Q #return angle Q angle = findq(3-1,2-4) ##store q in the variable angle a = 3-1 b = 2-4 angle=findq(a,b) print('The square of the angle is ',angle**2, 'and the angle is',angle) #now we can use this angle for something else
the angle Q is equal to -0.785 the angle Q is equal to -0.785 The square of the angle is 0.6168502750680849 and the angle is -0.7853981633974483
This is a more flexible function.
Review:
write a function that will use the general solution to the quadratic equation
\begin{equation} 0 = ax^2+bx+c \end{equation}Your function should do the following:
- take 3 numbers (a, b, c from the equation above) as inputs
- solve for the roots of the equation
- print the roots to the screen
- return the roots
Do this now.
Assignment 3: ass03 toy code
Do the ass03 assignment in your dropbox home.
- There is 1 notebook.
- There is 1 part in the notebook.
- Due date is written in the txt file in the ass03 directory in your dropbox home
The purpose of this assignment is to practice a little more using Jupyter and to learn (both for your instructor and for you) about how to use notebooks with assignments embedded in them, and to write a function.
Some more Jupyter knowledge (Markdown)
- Take this tutorial
https://www.markdowntutorial.com/
- jupyter's help on markdown cells
- book on LATEX math formatting
practice making headers, lists, and math equations in juytper using the markdown cell type
- Make the is list in Jupyter
- typeset the following equation in Jupyter
User input
Sometimes when you need input from the user the easiest is to use input.
Here is an example.
ui = input('type something ')
print(ui)
ui+1
ui = float(ui)
print(ui+2)
Control Structures: if
Now to Controlling things
Computers are not useful unless they make our lives better or easier. Consider the following situations:
- adding up all of numbers from 1:100
- adding up all of the even numbers from 1:100
- determining the square root of a list of numbers
- reading a file line by line until you reach the end of the file
Here is the first one in code form:
x = 1+2+3+4+5+6+7+8+9+10+11+12+13+14+15+16+17+18+19+20+21+22+23+24+25+26+27+28+29+30+31+32+33+34+35+36+37+38+39+40+41+42+43+44+45+46+47+48+49+50+51+52+53+54+55+56+57+58+59+60+61+62+63+64+65+66+67+68+69+70+71+72+73+74+75+76+77+78+79+80+81+82+83+84+85+86+87+88+89+90+91+92+93+94+95+96+97+98+99+100 print(x)
5050
Typing all of this is annoying, and if we have to type everything then the computer has not made our lives easier, nor have we gained any time if we want to repeat this computation, or alter it (only the odd numbers), or expand it 1 to 1000.
So most computer languages contain structures that are time saving devices or allow for the program to alter its computational trajectory if a situation changes.
In python there are 3 that are important for this course.
- if then else
- for loops
- while loops
To help you understand how theses structures work, we will use children's games.
If then else
This control structure can be understood under the context of the children's game Simon Says
Rules of the game:
if (the leader says "Simon say's" before they tell you to do something): [condition] then you do it [conditional action] else: don't do it [default action] end of turn
Translation of this previous block into code,
the text in the ( ) after the if statement is a testable condition (usually a mathematical comparison).
If this condition is True (sometimes designated as 1) then the [conditional action] code block is performed.
The else tells the computer that the default code block is coming up next
If the condition is False (sometimes designated as 0) then the default code block is performed.
The end lets the computer know that the control structure is over
Here is a toy example of using the if statement to see if two numbers are equal
x = 4 if x == 3: #in python the conditional if line ends in a colon (:) print('x was 3') x = x**2 print('now it is nine') else: # there is also a colon after else: print('x was not 3') # in python a code block (the whole if block) is "ended" by returning back to the # same column number as the i in if print('Thanks for playing!')
x was not 3 Thanks for playing!
NOTE:
- the = symbol is for setting something equal
- the == symbol is for comparison
There are many kinds of testable conditions that can go after the if statement, but some common ones are
- == (is equal to) Note that there are 2 equal signs next to each other. This is for comparison.
- < (is less than)
- <= (is less than or equal to)
- >= ( is greater than or equal to)
- > (is greater than)
- != not equal to
Here is another kind of if testing:
refletters = 'abcdefg' #this is a string questionletter = 'g' #here is another one if questionletter in refletters: #remember to end with a : print(questionletter, 'was in the set',refletters) else: #a new code block begins so remember to end with a : print(questionletter, 'was NOT in the set',refletters) print('Thanks for playing')
g was in the set abcdefg Thanks for playing
NOTE:
- the in conditional test means just what it sounds like "in"
- Also try if questionletter not in refletters and see what that does.
- Also try if questionletter == refletters and see what that does.
- <
- >
- ~=
These logical notations work as well but for strings may not be as easy to understand. The in and not use in these conditional tests are examples of how Python was designed to be easy to learn: Python was born to be readable.
else if or elif
There is also the possibility of performing multiple comparisons within one if then else structure. To do this use else if or elif in python. Here is an example:
#this program tests if the number stored in x to be greater than 3 # and if it is then set it to a maximun value of 3 x = 3.9 if float(x) == 3: # the colon tells python that a code block has begun print('x was 3') elif x > 3: # another code block so you need another colon at the end of the line print('x was greater than 3') x = 3.0 ## redefine x as 3 elif x < 3: # another code block so you need another colon at the end of the line print('x was less than 3. So I letter it pass.') else: # another code block so you need another colon at the end of the line pass #this command is only requred if this else or elif choise is empty msg = 'x is now {:.1f}'.format(x) ## some pretty printing of the result print(msg)
x was greater than 3 x is now 3.0
- what happens when you change what x = equal to? Try the values 4.5, -12 and 0
- what happens when you set x = 'v' (a string)?
- this raises an error and crashes the program
- we will learn to deal with this later
- Why? you can not compare a number (floating point or integer) and a string
- Comparisons need to be made between the same types of python objects (type whos in a cell to get a list of variables and their types)
Boolean logical combinations
Another way to combine multiple conditional comparisons is to combine two comparisons with Boolean expressions or operators
Extra reading on Boole and his Logic
Some example ideas:
- The 3 basic Boolean operators are
- and
- or
- not
- If you speak English then you know what these mean
- and both sides of the and needs to be true for the whole statement is true
- or either side of the or needs to be true for the whole statement is true
- not is often called an inverter. Its presence in a statement inverts the condition
- True becomes False
- False becomes True
- The same ideas apply to programming languages
- Here is an example:
#this program checks to see if x to be greater than 3 and less than 1 x = 5 if x <= 3 & x >= 1: #both must be true to for the conditional action to be executed print('x was between 1 and 3') #conditional action else: print('x was outside of the range 1 and 3') msg = 'x is {:.1f}'.format(x) print(msg)
x was outside of the range 1 and 3 x is 5.0
NOTE: the syntax (symbols and commands) used in Python for Boolean operators are either the simple English words (and, or , not) or 3 symbols:
| Boolean operator | word for symbol | symbol |
|---|---|---|
| and | ampersand | & |
| not | bang | ! |
| or | pipe or vertical bar | https://en.wikipedia.org/wiki/Vertical_bar |
General form
Note that the elif and else parts are optional, but make your code cooler looking.
if condition: code block elif other condition: #optonal code block #optional else: #optional defaultcode block # optional
Learning Review:
Write a function that
- takes one input
- determines if the number is greater than 0, less than zero, or zero.
- and prints the string '+', 'nil', or '-' depending upon the conditional testing results
Learning Review: Answer
def numtest(x): if x < 0: msg = '-' elif x > 0: msg = '+' elif x == 0: msg = 'nil' print(msg) return ## test the function numtest(75)
+
Strings, Lists
Strings
Before we jump into lists and arrays let us start with strings. We have already seen some string usage in this tutorial, but some ideas about lists can be learned from using stings.
First a string is surrounded by either single or double quotes. They both work, but you must used either single or double to define one string. Don't mix them! A string contains zero, one or more characters. For example:
word = 'python'
The variable word contains a string with 6 characters in it.
There are some special characters that can be in strings.
- \n results in a newline
- \t results in a tab
- \' puts a single quote mark in the string
Here is some usage
words = 'This is a string in \t \t python. \n\t Isn\'t it awesome!' print(words)
This is a string in python. Isn't it awesome!
Sometimes you want to put special characters in the string as normal characters. Here is an example and a solution
- use the r before the string definition to specify that it is raw and not interpreted.
path = 'C:\some\name_of_directory' print(path) betterpath = r'C:\some\name_of_directory' print(betterpath)
C:\some ame_of_directory C:\some\name_of_directory
Concatenation
Strings can be concatenated.
What is Concatenation?
- The linking of things together in a chain or series
- Strings are concatenated with the + between strings
Here are some examples:
## un un un + ium name = 3 * 'un' + 'ium' # NOTICE the + between the un and ium + means concatenate in strings print(name)
unununium
NOTE: notice the 3* 'un' in the code above. Some operators (multiply) can be used on strings. These operations mean to take 3 of the strings in a row, or concatenate 3 of the strings. Try using other operators on strings (divide, addition, subtraction), you will notice an error.
indexing and strings
From the python tutorial: https://docs.python.org/3/tutorial/introduction.html#strings
Strings can be indexed (subscripted), with the first character having index 0. (Some programming languages have a separate character type of variable. In Python there is no separate character type; a character is simply a string of size one. If you don't understand this set of statements, you can ignore it.)
Here are a bunch of examples of indexing strings:
word = 'Python' print(word[0]) # print the character in position 0 print(word[5]) # print character in position 5 ##extraction of subsets of the string x = word[1] #extract the second letter in the string or charactor as index = 1 print('the string\'s second letter or index 1 is',x) x = word[-1] # extract the last character in the string print('last is ',x) x = word[-2] # extract the second-last character in the string print('second to last is',x)
P n the string's second letter or index 1 is y last is n second to last is o
NOTE: Since -0 is the same as 0, negative indices start from -1. NOTE: Indices can have positive and negative values (see the last to examples in the code block above.)
slicing stings
In addition to indexing, slicing is also supported. While indexing is used to obtain individual characters, slicing allows you to obtain a substring:
word = 'Python' ### substrings x = word[0:2] # characters from position 0 (included) up to 2 (excluded) print('substring',x) x = word[2:5] # characters from position 2 (included) up to 5 (excluded) print('substring',x) print(word[:3] + word[3:]) ## see note below
substring Py substring tho Python
Note how the start is always included, and the end always excluded. This makes sure that s[:i] + s[i:] is always equal to s. This behavior is the main reason why python starts its indexing at zero. Here are some more demonstrations:
word = 'Python' ### substrings concatenation x = word[:2] + word[2:] print('notice that ',x,' is the same as ',word) x = word[:4] + word[4:] print('Also, notice that ',x,' is the same as ',word)
notice that Python is the same as Python Also, notice that Python is the same as Python
Slice indices have useful defaults;
- an omitted first index defaults to zero,
- an omitted second index defaults to the size of the string being sliced or the end of the string.
word = 'Python' ### substring slicing x = word[:2] # characters from the beginning to position 2 (excluded) print('characters from the beginning to position 2 (excluded) are => ',x) x = word[4:] # characters from position 4 (included) to the end print('characters from position 4 (included) to the end are => ',x) x = word[-2:] # characters from the second-to-last (included) to the end print('characters from the second to last (included) to the end are => ',x)
characters from the beginning to position 2 (excluded) are => Py characters from position 4 (included) to the end are => on characters from the second to last (included) to the end are => on
One way to remember how slices work is to think of the indices as pointing between characters, with the left edge of the first character numbered 0. Then the right edge of the last character of a string of n characters has index n, for example:
+---+---+---+---+---+---+ | P | y | t | h | o | n | +---+---+---+---+---+---+ i 0 1 2 3 4 5 6 j -6 -5 -4 -3 -2 -1
- The first row of numbers gives the position of the indices 0…6 in the string;
- The second row gives the corresponding negative indices.
The slice from i to j consists of all characters between the edges labeled i and j, respectively.
Note: For non-negative indices, the length of a slice is the difference of the indices, if both are within bounds. For example, the length of word[1:3] is 2.
Attempting to use an index that is too large will result in an error. Type and see what happens:
word = 'Python' word[42]
In contrast: out of range slice indexes are handled gracefully when used for slicing:
word = 'Python' x = word[4:42] print(x) print(word[42:]) ##what is this doing prints an empty string
on
Learning Review
Using this template above: Figure out how to extract some of these examples from the word variable (word = 'Python'):
- hon
- Pyth
- Pytho
- n
- tho
- ytho
Learning Review: Answers
word = 'Python' print(word[3:]) print(word[:4]) print(word[:5]) print(word[-1]) print(word[2:5]) print(word[1:5])
hon Pyth Pytho n tho ytho
the immutable string
Python strings cannot be changed — the computer science word for this is immutable. Therefore, assigning to an indexed position in the string results in an error: Here is a definition: https://docs.python.org/3/glossary.html#term-immutable
Here is an example in which we need to lower case the first letter of Python.
word = 'Python' #define word the first time no problem word[0] = 'p' # try to change it you will get an error message word = 'python' # redining the whole variable is no problem
Here is another way to fix this particular problem (see above).
word = 'Python' print('p' + word[1:])
python
determine the length of a string
use the len() function to get the length of things:
word = 'supercalifragilisticexpialidocious' length = len(word) print('the word ', word,' has ',length,' characters.')
the word supercalifragilisticexpialidocious has 34 characters.
Learning Review A:
Write a function that compares the lengths of two strings and returns the longest word with its length.
- function
- 2 inputs
- compute length of each string
- compare lengths
- output string and its length (HINT: last line of the function should include return string, length )
Learning Review B: More Advanced (to do at your own pace)
Write a function that compares the lengths of two strings and returns the longest word with its length.
- function
- 3 inputs
- compute length of each string
- compare lengths
- output string and its length (HINT: last line should include return string, length )
lists
From Python tutorial https://docs.python.org/3/tutorial/introduction.html#lists
Python knows a number of compound data types, used to group together items. The most versatile is the list, which can be written as a list of comma-separated items (numbers, strings, other lists, whatever) between square brackets [ ] . Lists might contain items of different types, but usually for this tutorial the items all have the same type.
Defining lists
squares = [1, 4, 9, 16, 25] ## define a list by enclosing csv in [] print(squares) letters = ['A','B','C','D'] print(letters)
[1, 4, 9, 16, 25] ['A', 'B', 'C', 'D']
indexing and slicing lists
Just like strings in a previous section (and all other built-in sequence types), lists can be indexed and sliced:
squares = [1, 4, 9, 16, 25] ## define a list by enclosing csv in [] x = squares[0] # indexing returns the item print('x =', x) y = squares[-1] print('y =', y) z = squares[-3:] # slicing returns a new list print('z = ',z)
x = 1 y = 25 z = [9, 16, 25]
All slice operations return a new list containing the requested elements. This means that the following slice returns a new copy of the list:
squares = [1, 4, 9, 16, 25] ## define a list by enclosing csv in [] x = squares[:] print('x = ',x)
x = [1, 4, 9, 16, 25]
Concatenation of lists
Lists also support operations like concatenation:
squares = [1, 4, 9, 16, 25] ## define a list by enclosing csv in [] x = squares + [36, 49, 64, 81, 100] # concatentation with + print('now x = ',x)
now x = [1, 4, 9, 16, 25, 36, 49, 64, 81, 100]
The multiply operator on a list behaves like strings. Here is a demonstration:
squares = [1, 4, 9, 16, 25] ## define a list by enclosing csv in [] x = 3* squares # repeated concatentation with * print('now x = ',x)
now x = [1, 4, 9, 16, 25, 1, 4, 9, 16, 25, 1, 4, 9, 16, 25]
NOTE: this behavior of lists always seems weird to me. It is not what I would want to happen. We will learn later several ways to operate on all of the elements in a sequence (List, arrays, dataframes).
Lists are Mutable
Unlike strings, which are immutable, lists are a mutable type, i.e. it is possible to change their content:
squares = [1, 4, 9, 16, 25] ## define a list by enclosing csv in [] cubes = [1, 8, 27, 65, 125] # something's wrong here print('something is wrong here: ',cubes) print('but ',4 ** 3, ' is the cube of 4') # the cube of 4 is 64, not 65! cubes[3] = 64 # replace the wrong value print('now cubes = ',cubes)
something is wrong here: [1, 8, 27, 65, 125] but 64 is the cube of 4 now cubes = [1, 8, 27, 64, 125]
Adding items to the end of a list
You can also add new items at the end of the list, by using the append() method (we will see more about methods later)
squares = [1, 4, 9, 16, 25] ## define a list by enclosing csv in [] cubes = [1, 8, 27, 65, 125] # something's wrong here print('cubes was initially defined as ',cubes) cubes.append(216) # add the cube of 6 cubes.append(7 ** 3) # and the cube of 7 print('appending 2 new values to cubes. now cubes is ',cubes)
cubes was initially defined as [1, 8, 27, 65, 125] appending 2 new values to cubes. now cubes is [1, 8, 27, 65, 125, 216, 343]
Assignment to slices is also possible, and this can even change the size of the list or clear it entirely:
letters = ['a', 'b', 'c', 'd', 'e', 'f', 'g'] print('letters is ',letters) # replace some values letters[2:5] = ['C', 'D', 'E'] print('replacing indices 2-5 yields',letters) # now remove them letters[2:5] = [] print('droping the indeces 2-5 yeilds',letters) # clear the list by replacing all the elements with an empty list letters[:] = [] print('clearing the whole list yeilds and empty list',letters)
letters is ['a', 'b', 'c', 'd', 'e', 'f', 'g'] replacing indices 2-5 yields ['a', 'b', 'C', 'D', 'E', 'f', 'g'] droping the indeces 2-5 yeilds ['a', 'b', 'f', 'g'] clearing the whole list yeilds and empty list []
Getting the length of a list
The built-in function len() also applies to lists:
letters = ['a', 'b', 'c', 'd'] length = len(letters) print('The length of this list is ',length)
The length of this list is 4
Lists of Lists
It is possible to nest lists (create lists containing other lists), for example:
### hand coding a list of lists LL = [['I','think', 'we'],['are','here']] print('hand coded list of lists ',LL) ## or more programatically a = ['a', 'b', 'c'] #make 1 list n = [1, 2, 3] #make another list x = [a, n] # now make a list of lists or nested lists print('look at what x is ',x) ##indexing a nested list b = x[0] print('the 0th index in x is ',b, ' and is the sublist a =>',a) ## indexing a nested list deepinside = x[0][1] ## note the 2 sets of square brackets print('here is an item within the 0 index nested list = >',deepinside) print('here is an item within the 1 index nested list = >',x[1][2])
hand coded list of lists [['I', 'think', 'we'], ['are', 'here']] look at what x is [['a', 'b', 'c'], [1, 2, 3]] the 0th index in x is ['a', 'b', 'c'] and is the sublist a => ['a', 'b', 'c'] here is an item within the 0 index nested list = > b here is an item within the 1 index nested list = > 3
learning Review
Create a funny sentence function. This function will pick 3 random indices and pull words from 3 random lists, then print to the screen the concatenated funny sentence.
Example: list of subjects: I, you, donkey list of verbs: eat, like, poop list of objects: bread, corn, hippos
randomly choose three indices 0,1,2 then your funny sentence that should be printed to the screen is:
I like hippos
Use the numpy function randint to get an integer that can be used as an index.
import numpy as np #this command loads all of the functions in numpy and labels them np from numpy.random import randint ## this command loads the function randint in to memory for later use subjects = ['I', 'You'] idx = randint(0,high=2) # the 0 is the lowest integer (0th index) # ths high=3 is one count above the highest integer that you need # to make this flexible maybe don't hard code the high = 2 # but get it from the list print('idx is ',idx)
idx is 1
Assignment 6 (ass6) Part 1: create a function that
- takes 2 inputs
- 1\(^{st}\) is a list of floating point numbers (1.0, 3.12, etc.)
- 2\(^{nd}\) is an index that is less than the length of the list
- takes the number in the list at the index position and replaces it with 13
- appends the number 114 to the end of the list
- uses the built-in function sum() to sum all of the items in the list
- returns the sum of the newlist
Due Date is in the directory
Assignment 6 (ass6): Part 2 List slicing (10 points)
create a function that
- takes 1 input
- the input is a list of items 'dog',1,'R',1.0, 3.12, or whatever with length > 5
- checks to see if the list is > 5 items long
- prints an error message if the input list is too short, and returns a 1
- otherwise takes a slice of the 1,2,4 index of the list
- returns the returns the slice of the list
Due Date is in the directory
control structures: for loop
This control structure is a little difficult to compare to a children's game. The most similar game is a relay race. Your goal in this race is to carry a stick around the track a specific number of times.
for each time around the track carry the baton when you have looped around the track the required amount of times the race is at an end
Use a for loop to do a task a specific number to times. Here is an example
for i in [1,2,3,4,5,6,7,8,9,10]: print(i)
1 2 3 4 5 6 7 8 9 10
In the previous example, [1,2,3,4,5,6,7,8,9,10] is a list. I contains a list of integer numbers.
You would read this code block as "for i equals items in the list", [1,2,3,4,5,6,7,8,9,10], then display i then end the code block. In a for loop, the iterator variable i in the case above takes on the values from the list 1, 2, 3, 4, 5, ….
range()
What would be really nice is if we had a way to generating this list [1,2,3,4,5,6,7,8,9,10] rather than typing it. Use the command range()
This example should give the same output.
for i in range(10): print(i)
0 1 2 3 4 5 6 7 8 9
The range() function can take up to 3 inputs.
- range(begin,end,stepsize)
Let's use this range() function to add up some numbers
start = 1 stop = 101 step = 1 count = 0 ## this is the seed of the count for i in range(start,stop,step): count = count + i print(i, count)
1 1 2 3 3 6 4 10 5 15 6 21 7 28 8 36 9 45 10 55 11 66 12 78 13 91 14 105 15 120 16 136 17 153 18 171 19 190 20 210 21 231 22 253 23 276 24 300 25 325 26 351 27 378 28 406 29 435 30 465 31 496 32 528 33 561 34 595 35 630 36 666 37 703 38 741 39 780 40 820 41 861 42 903 43 946 44 990 45 1035 46 1081 47 1128 48 1176 49 1225 50 1275 51 1326 52 1378 53 1431 54 1485 55 1540 56 1596 57 1653 58 1711 59 1770 60 1830 61 1891 62 1953 63 2016 64 2080 65 2145 66 2211 67 2278 68 2346 69 2415 70 2485 71 2556 72 2628 73 2701 74 2775 75 2850 76 2926 77 3003 78 3081 79 3160 80 3240 81 3321 82 3403 83 3486 84 3570 85 3655 86 3741 87 3828 88 3916 89 4005 90 4095 91 4186 92 4278 93 4371 94 4465 95 4560 96 4656 97 4753 98 4851 99 4950 100 5050
Here is another way to do the same thing, but it stores the values of i in a list:
start = 1 stop = 101 step = 1 L = [] ## this starts a list to put the numbers in for i in range(start,stop,step): L.append(i) print(i, L) Ls = sum(L) print('The sum is ', Ls)
100 [1, 2, 3, 4, 5, 6, 7, 8, 9, 10, 11, 12, 13, 14, 15, 16, 17, 18, 19, 20, 21, 22, 23, 24, 25, 26, 27, 28, 29, 30, 31, 32, 33, 34, 35, 36, 37, 38, 39, 40, 41, 42, 43, 44, 45, 46, 47, 48, 49, 50, 51, 52, 53, 54, 55, 56, 57, 58, 59, 60, 61, 62, 63, 64, 65, 66, 67, 68, 69, 70, 71, 72, 73, 74, 75, 76, 77, 78, 79, 80, 81, 82, 83, 84, 85, 86, 87, 88, 89, 90, 91, 92, 93, 94, 95, 96, 97, 98, 99, 100] The sum is 5050
NOTE: the command sum(). It does what you think it should. It sums the list L.
List processing with for loops with range()
This list generator (range()) is good for generating indices for accessing components in a list or other sequence data structure.
L = ['The','cat', 'has','a','scary','smell'] for i in range(len(L)): print(L[i])
The cat has a scary smell
In Python, this iterating over a list is very powerful, especially because the list can contain any python structure.
L = ['The','cat', 'has','a','scary','smell'] for word in L: print(word) #Notice how there is no math here at all
We can also access more than one list by using an indices generator.
L = ['The','cat', 'has','a','scary','smell'] P = ['THE','CAT', 'HAS','A','SCARY','SMELL'] for i,word in enumerate(L): ##use the enumerate around a list to create an index generator print(i,word,P[i]) # i is the index of the list and word is the item in the list
0 The THE 1 cat CAT 2 has HAS 3 a A 4 scary SCARY 5 smell SMELL
Learning Review: add up some stuff
Create a function to add up all of the numbers in a range
- use a for loop to step through the range of numbers
- use range: 12 to 115
- also figure out a quick change to your range of every 3rd number
- print and returns the sum
Learning Review: Answer
# learning review check add up some stuff #first do it not in a function start = 12 stop = 115 step = 1 count = 0 for i in range(start, stop + 1, step): count = count + i print(i,count) ## then make a function def addmup(start,stop, step): #put the parts defined by the user as inputs ''' Adds up all the numbers from start to stop step is the step size between the start and stop, 1,2,3, etc. ''' count = 0 ## you start adding at zero for i in range(start, stop + 1, step): ## notice stop + 1, range before the value in stop count = count + i print('The sum of all values in the range (',start,' to ',stop,') is',count) return count ## test out your function addmup(12,115,1)
115 6604 The sum of all values in the range ( 12 to 115 ) is 6604
Learning Review: Fibonacci list
write a function that
- takes 1 input as an integer (n)
- creates a list that contains the Fibonacci sequence up to and including the nth
- returns the list
Learning Review: Fibonacci list: Answer
def fibonacci(n): ''' function to generate a list of the first nth numbers in the Fibonacci sequence n is the integer number in the fibonacci sequence n must be larger than 1 ''' F1 = 1 #first number in the Fibonacci sequence F2 = 1 # second number in the Fibonacci sequence fiblist = [F1,F2] # seed the list with the first two fibonacci numbers for i in range(1,n-1): # start loop up to n, note: start with 1 because that is how most people count, n-1 already have fisrt 2 in list, remember in range(stop) the stop number is 1< stop F = F1+ F2 # calculate the next number fiblist.append(F) #store the F in running list F1= F2 # Shift current F1 to F2 F2 = F # shift current F to F1 return fiblist #return the list print(fibonacci(12))
[1, 1, 2, 3, 5, 8, 13, 21, 34, 55, 89, 144]
Check if an item is in a list
if 'cat' in ['cat','dog','cow']: print('yes')
yes
ass7: Fibonacci 1
Write a function that will
- take 1 input that is a positive integer (n)
- if the input is not an integer return an error of 1 (HINT: use the isinstance function link)
- and print an error message
- calculate the nth Fibonacci number
- return the Fibonacci number
ass7: Fibonacci 2
write a function that
- takes an integer as an input
- checks if the input is a Fibonacci number
- returns a True if it is a Fibonacci number
- returns a False if it is not a Fibonacci number
ass7: Fibonacci 3
Format in Markdown in the cell provided the Equation:
\begin{equation} F_n = F_{n-1} + F_{n-2} \end{equation}- Due Date Feb 16, 2018 at noon
Arrays and Matrices for data storage
As we learned in the last several modules, lists are powerful Python objects that can be used to store information or data. However there are several (2) limitations with lists that make them less useful for storing data.
Specifically,
- mathematical operations on a list don't work in a useful way.
- ???
Here is an example this limitation: Let's say you were needed to calculate the square of a list of numbers (such as in the calculation of the standard deviation.)
Try this in python
data = [15.2, 16.7, 17.34, 14.6, 13.9] data*data # this is the calculation for the square of a number data ** 2 # nor can you do this
Notice how it gives an error. You can not multiply 2 lists together nor can you take the power 2 of a list.
Here are 2 ways of doing this operation using a for loop Note there are more variables used here than are necessary. I am using them to make the example more understandable.
data = [15.2, 16.7, 17.3, 14.6, 13.9] ## the org data sqdata = [ ] ## create an empty list to store the squares for d in data: ## us 'in' to get the items out of the data list one at a time sqd = d**2 ## square the data and store it in a new variable sqdata.append(sqd) ## append the square to the list print(data) print(sqdata)
[15.2, 16.7, 17.3, 14.6, 13.9] [231.04, 278.89, 299.29, 213.16, 193.21]
data = [15.2, 16.7, 17.3, 14.6, 13.9] ## the org data sqdata = [ ] ## create an empty list to store the squares for i in range(len(data)): ## create an increasing index 'i' to access the items in data d = data[i] ## extract the item from list 'data' sqd = d*d ## square the data and store it in a new variable sqdata.append(sqd) ## append the square to the list print(data) print(sqdata)
[15.2, 16.7, 17.3, 14.6, 13.9] [231.04, 278.89, 299.29, 213.16, 193.21]
Using a for loop to process lists can be tedious. If you have a huge list it is also very slow. (This is a fact that is not self-evident from what we have covered so far. You will have to either accept this statement or you can read about it link.
So the Python community have created additional Python object types to help with this kind of problem. They are array, matrix, and dataframe. These are all very similar to List, but they have special properties that are useful for data analysis.
You might envision arrays as a kind of data table or spreadsheet-type of computer object.
Look at this table.
| 1 |
| 2 |
| 3 |
| 4 |
It is basically an array of dimensions 4 by 1 (4x1) or (4,1). [ Notice the dimensions are (row, column). ] And here
| 1 | 2 | 3 | 4 |
is an array of dimensions 1 by 4 (1x4) or (4,1). These are both considered 1D arrays, which means that they only have 1 dimension that is greater than 1.
Now consider the following array
| 1 | 2 |
| 3 | 4 |
What are its dimensions?
- ANSWER: (2,2) two by two
The only limitation with thinking of arrays as a kind of table is that most people don't think of data tables as having more than 2 dimensions. But arrays can have as many dimensions as you can imagine (or whatever the limit of the computer language Creator's imagination was).
Here is some code for understanding arrays.
creating arrays
To create/operate/use arrays we must use the numpy module.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np A = np.array([4,5,6,7]) #turn a list into an array print(type(A)) ## print the type of object print('the shape of the array is ',A.shape) # print the shape of the array.
<class 'numpy.ndarray'> the shape of the array is (4,)
Type %whos in the cell of Jupyter (the % means magic command in Jupyter) and you should see in the list:
Variable Type Data/Info ------------------------------- A ndarray 4: 4 elems, type `int64`, 32 bytes np module <module 'numpy' from '/ho<...>kages/numpy/__init__.py'>
Notice the variable A type is ndarray. You can also access this information with the type(function). ndarray means n-dimensional array.
Use the .shape method to extract the shape of an array. Why does it say (4,) instead of (4,1)? This is just Python's way of saving space/memory.
Element-wise Operations on an array
The reason of using arrays is to operate on them easily. Let's take the problem of squaring a list above. Here is how to solve this problem with numpy arrays.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np datalist = [15.2, 16.7, 17.3, 14.6, 13.9] ## the org data data = np.array(datalist) ## first make the list into an array sqdata = data**2 ## square the data print(datalist) print(sqdata) print(type(sqdata))
[15.2, 16.7, 17.3, 14.6, 13.9] [231.04 278.89 299.29 213.16 193.21] <class 'numpy.ndarray'>
Whenever you use an operator on an np.array object it operates on each element and returns a new array. Try a several
- data-sqdata
- data*sqdata
- data + data
These all do element-wise calculations.
indexing and slicing
Array indexing and slicing works just like a list.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] ) # onestep create of array print('The zeroth index of the array is ', data[0]) #indexing returns a number print('The first 3 items in the array are ', data[:3]) ## slicing returns an array
The zeroth index of the array is 15.2 The first 3 items in the array are [15.2 16.7 17.3]
modification of an array
*Array*s are like *list*s in that they can be modified.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] ) # onestep create of array print('The zeroth index of the array is ', data[0]) #indexing returns a number data[0] = 55.5 ## replace this zeroth indexed number print('The zeroth index of the array is now', data[0]) #indexing returns a number
The zeroth index of the array is 15.2 The zeroth index of the array is now 55.5
Special arrays
There are some special arrays that are easy to create with numpy.
- an array of all ones
- an array of all zeros
- an array of all random numbers
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np ## all ones A = np.ones( (3,3) ) ## the (3,3) is the shape that you want print('A is \n', A) ## the \n prints a new line print(A.shape) ## all zeros B = np.ones( (4,4) ) ## the (4,4) is the shape print('B is \n', B) ## the \n prints a new line print(B.shape) ## normally distributed random numbers center = 3 ##center of random numbers spread = 2 ## standard deivation of the random numbers shape = (2,4) ## shape of the array created. this is called a tuple C = np.random.normal(loc=center,scale=spread,size=shape) print('C is \n',C) ## the \n prints a new line print(C.shape) ## uniformly distributed random numbers D = np.random.rand(3,5) ## random numbers are spread out evenly between 0 and 1 print('D is \n ', D) ## the \n prints a new line print(D.shape)
A is [[1. 1. 1.] [1. 1. 1.] [1. 1. 1.]] (3, 3) B is [[1. 1. 1. 1.] [1. 1. 1. 1.] [1. 1. 1. 1.] [1. 1. 1. 1.]] (4, 4) C is [[2.73534909 3.65547042 0.63413359 0.98220572] [0.74696185 0.11992447 1.89381477 3.80934037]] (2, 4) D is [[0.86655031 0.5470224 0.66410465 0.57506126 0.35278717] [0.31438514 0.77715677 0.89651335 0.89601613 0.88361964] [0.13078894 0.40011216 0.58053045 0.14277761 0.59582208]] (3, 5)
Whole array manipulations (continued)
Another powerful utility of numpy is its built-in functions. Here are a few useful examples that do what you think they do.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] ) # onestep create of array print(data) print(np.sin( data )) print(np.sum(data)) print(np.median(data)) print(np.min(data)) ## determines the miniumn value in an array print(np.max(data)) ## determines the maximun value in an array print(np.std(data,ddof=1)) ## determine the standard deviation "ddof = 1" means use degree of freedom = 1, which is the words for n-1 in the bottom of the std. dev. equation.
[15.2 16.7 17.3 14.6 13.9] [ 0.48639869 -0.83714178 -0.99977443 0.89479117 0.9720075 ] 77.7 15.2 13.9 17.3 1.4258330898110059
numpy has about 99% of all of the functions that you might want. Here is a list in their reference page.
array methods
In the numpy module, the array objects have methods associated with them. A method is an operation on an object that can be applied to that object. Jupyter has several nice features that can help use understand some of these.
import numpy as np #this command loads all of the functions in numpy and labels them np data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] ) # onestep create of array print( np.sum(data) ) ## print the np.sum of the data ndarray print( data.sum() ) ## same as np.sum( ) but as a method: the dot .sum() makes it a method
77.7 77.7
The methods produce the same results as the functions. Note: Not all functions are methods. Try these methods:
- sum()
- mean()
- max()
- sin()
In Jupyter there are some convenience features. Type
import numpy as np #this command loads all of the functions in numpy and labels them np data = np.array([15.2, 16.7, 17.3, 14.6, 13.9] ) # onestep create of array # then type data. and depress the tab button, and wait, you will see a menu pop-up of all the methods
the Matrix
Python provides a special type of array called a matrix. The major differences between a matrix and an array is the manner in which mathematical operations are handled.
- Arrays review: math works element-wise
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np A = np.array([[ 2, 2.], [-2. , 2]] ) ## create array print(A*A) ## perform element-wise multiplication B = np.array([2,2]) print('The shape of B is', B.shape) ## notice how the shape is (2,) #you can also flip the rows and columns, this process is called transpose print('B is \n', B.T) # .T method for transpose print(B.T.shape) ## notice the shape now, transpose does not mean anything for a (2,) array
[[4. 4.] [4. 4.]] The shape of B is (2,) B is [2 2] (2,)
- Matrix: math works using the rules of Linear Algebra
We will learn some basics of linear algebra later in this tutorial, but for now here is an example of this linear algebra behavior of matrices. Unusual results from mathematical operation can result.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np A = np.matrix([[ 2, 2.], [-2. , 2]] ) ## create a matrix print(A*A) ## perform matrix multiplication, not element-wise multiplication B = np.matrix([2,2]) print('The shape of B is', B.shape) ## notice how the shape is (2,1) and not (2,) like an array #you can also flip the rows and columns, this process is called transpose print('B is \n', B.T) # .T method for transpose print(B.T.shape) ## notice the shape now # you can also convert between the array and the matrix data objects Am = np.matrix([[ 2, 2.], [-2. , 2]] ) print('Matrix type', type(Am)) Aa = np.array(Am) print('Now it is an array', type(Aa)) # the reverse is also true Aa = np.array([15.2, 16.7, 17.34, 14.6, 13.9]) print('Array data type', Aa) Am = np.matrix(Aa) print('Now it is a matrix', type(Am))
[[ 0. 8.] [-8. 0.]] The shape of B is (1, 2) B is [[2] [2]] (2, 1) Matrix type <class 'numpy.matrix'> Now it is an array <class 'numpy.ndarray'> Array data type [15.2 16.7 17.34 14.6 13.9 ] Now it is a matrix <class 'numpy.matrix'>
Some matrix specific operations are diag and trace. Here are some examples:
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np A = np.matrix([[ 2, 2.], [-2. , 2]] ) ## create a matrix print(A) print( np.diag(A) ) ## the diagonal of the matrix is an array D = np.diag(A) print( np.sum(D) ) ## sum of the diagonal is the trace # or T = np.trace(A) print(T)
[[ 2. 2.] [-2. 2.]] [2. 2.] 4.0 4.0
Assignment 8 Arrays (ass8)
create a function that will
- take as inputs 2 arrays of the same size
- element-wise add the two arrays together
- divide resulting sum-array by 2
- return the array return the resulting array
Assignment 9 Matrix math (ass9)
create a function that will
- take as inputs 2 matrices of the same size
- matrix multiply them together
- then calculates the trace for product of the two matrices
- returns the trace of the matrix multiplication (or the sum of the diagonal)
Dataframes (the named arrays)
Dictionaries adfad
Dictionaries are another kind of information/data storage object in python. We will not use them very much in this tutorial. They are a good segue into the next section dataframes, so that is why we are covering them here.
Python dictionaries are just like dictionaries in the real world. You put in a word and you get out a definition.
A dictionary is defined by either the word dict() or enclosing it in {}. Here are some examples.
Just like a list you can store anything in a dictionary.
import numpy as np #this command loads all of the functions in numpy and labels them np ## generic example: # name = {key:value,key:value} hr = {'Aaron':47, 'Ruth':60, 'Bonds':73, 'McGwire':70, 'Sosa':66 } print( hr['Bonds'] ) ### a simplified spectroscopic example wavelengths = np.array([200, 250, 300, 350, 400]) spec = np.array([0, 0.2, 0.5, 0.3, 0.001]) d = {'x':wavelengths, 'y':spec} y = d['y'] print(y)
73 [0. 0.2 0.5 0.3 0.001]
Dataframes
Python has another module called Pandas user guide. That was built to make certain tasks in python/numpy easier. In this module is defined a new data type object called a dataframe. You can think of these types as named arrays, or some kind of combination of array and dictionary.
Here is a continuation of the baseball example above:
import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd ## this command loads the pandas module hr = {'Aaron':47, 'Ruth':60, 'Bonds':73, 'McGwire':70, 'Sosa':66 } df = pd.DataFrame.from_dict(hr,orient='index',columns=['HR']) print(df) #the dataframe presents itself as a table with the rows labeled (called the index) and the column(s) labeled print('') print('****************** spacer *********************') print('Average Homeruns \n',df.mean()) ## you can operate on the dataframe, NOTE: the output of these operations is another dataframe print('****************** spacer *********************') #to get the values out of the dataframe use the .values method hrdf = df.mean() print(hrdf) print(hrdf.values) ## output as an array
HR
Aaron 47
Ruth 60
Bonds 73
McGwire 70
Sosa 66
****************** spacer *********************
Average Homeruns
HR 63.2
dtype: float64
****************** spacer *********************
HR 63.2
dtype: float64
[63.2]
Here is another example that is a little more relevant to chemistry
import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd ## this command loads the pandas module wavelengths = np.array([200, 250, 300, 350, 400]) ## wavelength labels for absorbance values spec1 = np.array([0, 0.2, 0.5, 0.3, 0.001]) ## absorbance values spec2 = np.array([0.2,0.5,0.1,0.0,0.0]) ## aborbance values d = {'spectrum 1':spec1, 'spectrum 2':spec2} ## make a dictionary df = pd.DataFrame.from_dict(d,orient='index',columns=wavelengths) ## make a dataframe print('Type for spec',type(spec1)) print('Type for d',type(d)) print('Type for df',type(df)) print('****************** spacer *********************') ## as a dataframe you can take the average spectrum print(df.mean()) print('****************** spacer *********************') ## or you can get the absorbacne values at a specific wavelength print(df[300])
Type for spec <class 'numpy.ndarray'> Type for d <class 'dict'> Type for df <class 'pandas.core.frame.DataFrame'> ****************** spacer ********************* 200 0.1000 250 0.3500 300 0.3000 350 0.1500 400 0.0005 dtype: float64 ****************** spacer ********************* spectrum 1 0.5 spectrum 2 0.1 Name: 300, dtype: float64
CANCELED Activity Assignment 10 (ass10) (maybe move to loading data section)
Making Graphics and Figures
All discussion of spectroscopic (and most other) data is often accompanied by some kind of graphical representation of the data. This section will describe the fundamentals of graphical representation of data plus the graphical representation of the meaning of the data. \footnote{I don't actually know what this last clause means, but it sounds cool.}
It is my philosophy that all graphical data presentations must stand by themselves, without oral explanation. So when you hand a printout of data to your colleague or boss, you should not have to explain every detail of the graphic. Some obvious parts of the presentation such as axes labels, legends, and captions easily enhance the understanding of the graphic. Herein are a set of examples from which you can learn.
There are several types of data presentations, and the choice of which to use depends upon the type of data and the intended usage of the data presentation.
There are 6 basic types of data presentation graphs that are commonly used by chemists: line graph, scatter plot, histograms, contour plots, surface plots, and images. (I guess there is also a bar graph and pie graph, but I don't really know when to use these types.) There are many perturbations within each of these graph types including combinations of each within one presentation.
There are several intended uses for these data presentations, and these uses in part dictate how the figures are created.
- Computer-based inspection: For me this is the most useful, because I can interact with the data (zoom in on specific parts, inspect the peak max or min, overlay 2 spectra, etc.).
- Paper-based lab notebook: This printed version does not require super high resolution; I will refer to this type of presentation as a thumbnail print.
- Publication Quality: Related to the thumbnail print is the publication quality print, which is a higher resolution file size and may also have additional annotations and figure captions.
- Beamer Presentation: Also related to the printed version is for inclusion into a beamer presentation or image projection system for giving talks, which usually don't require a high resolution graphic format.
All of these can be created with python, numpy, matplotlib, pandas, and scipy. Additionally, these data presentations may need to be annotated to enhance the understanding of the information presented.
In the following subsections are examples of these different types of graphics and their uses.
General Philosophy for what type of graph to use
In the introduction above were listed 6 types of graphs: line graph, scatter plot, histograms, contour plots, surface plots, and images. So a question might arise in your mind, "Which type do I use?" There have been many discussions over the years about this question. Here are a few links:
- https://www.edwardtufte.com/tufte/
- https://journals.plos.org/ploscompbiol/article?id=10.1371/journal.pcbi.1003833
- http://mirrors.ctan.org/graphics/pgf/base/doc/pgfmanual.pdf
This is an unsettled issue that is full of opinions. Herein I will present my opinions as well as some rules that are settled by conventions within the field of chemistry (mostly IUPAC). I will delineate between my opinions and conventions, by stating when something is my opinion. All other statements should be assumed to be based upon best practices (conventions).
These next 6 subsections are all my opinion, but I believe that most scientist would agree with them.
when to use a scatter plot
Whenever you need to fit a smallish set of data with a mathematical function, it is better to plot the data with markers, and the fit with a line. My reasons are that this formula produces a clear distinction between the actual data and the mathematical function. You don't want the viewers of you data presentation to think that you have more data than you do.
Some examples are shown in the Figure below.
Figure 12: Example scatter plots with the data plotted with markers and the fit with a line
When to use a line plot
Whenever you have loads of data points that are follow a predictable trajectory, like in a spectrum or chromatorgram, using markers will make the line so thick that you might inadvertently hide the curvature. In these situations, I believe you can use a smoothly curving line to present the spectrum, chromatorgram, etc. Some examples are shown in the figure below.
Figure 13: Example line plots with lots of points
When to use histograms
Whenever you have 1D data that is from a set of repeated measurement, and your goal is to represent the spread (how repeatable) in the data, then a good choice is a histogram. An example, might be if you were monitoring the arsenic concentration level in bunch of samples.
Figure 14: Example of a histogram of repeated measurements
When to use contour and surface plots and images
In chemistry data, these 3 types of plots can be used interchangeably. One exception is when presenting data that is in fact an image (e.g. photograph, SEM image, etc.). The types of data that utilize these types of graphs fall in to 2 categories:
- correlated/anticorrelated data
These types of data are commonly found in 2D correlation spectroscopy. Some of you may have learned about 2D NMR measurements such as COSY ( link ) and HETCOR. These data are often presented as contour plots. These NMR measurements are a subset of experiments that are part of the generalized 2D spectroscopy type of measurements. Here is a very good article link.
In all of these cases, the X and Y dimensions are some kind of spectrum (X and Y can be the same or different spectra or even spectral types), and the in the middle are the relative heights of the correlations between the two (X and Y) spectra portrayed as contours.
Here is an example of a contour plot of the correlation of UV spectra during a chemical reaction.
Figure 15: Example of generalized 2D spectroscopy
This same kind of data could also be displayed as a 3D surface. In fact a contour plot is just a flattened surface plot. Here are the same data shown in 3 different graphs. Notice how the contour graph and image really need a color bar to the right so that the reader can know what the colors represent.
Figure 16: Examples of a contour, surface plot and an image of the same data
- data in which multiple two axes are the same kind of data (spatial) and the 3rd is a different kind of data (concentration, spectra band height, etc.)
You have seen this kind of data before in the following figure.
Figure 17: Finger print, DNA, RDX
Generic Example
Each of the examples within this section of the tutorial will contain several parts. Not all of the parts will be used in every example, but the part is listed to keep the examples consistent. Here is a generic list of the parts:
1. import modules (you only need this part once per jupyter notebook) 2 get or make the data 3. process the data 4. plot the data 5. plot the data 6. save the data in the required format 7. house keeping (this is part of the tutorial but not for you to copy)
plotting useless data
Let us start by just learning to plot using made-up, simulated data. There are 3 kinds of graphs that are common in chemistry:
- a continuous graph of data without markers
- an xy scatter graph with markers
- a histogram of data from repeated measurements with some spread in the data
simulated spectrum
For spectra, chromatograms, and other continuous data it is better to plot the data with a line without markers. The main reason is that the markers would be so close together that they would interfere with the viewer's perspective of the shape of the curve. The default for the matplotlib library is to add space to either sided of these types of graphs. Personally, I don't like this. For now we will leave then there, and learn how to eliminate them latter in this tutorial.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions # you only need to load the modules once per file import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get/make the data x = np.arange(200,400) # these will be the x values in the graph for a simulated UV spectrum xc = 250 # this is the center of the band sigma = 10 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the lambda max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve (idea shape of lots of data) ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ## ax is the variable for the axis object where you want to work ax.plot(x,spec) ## .plot(x,y) method to make a plot on the ax object ax.set_xlim((200,400)) ## set x dim edges of the plot ax.set_xlabel('Wavelength /nm') ## LaTeXworks here ax.set_ylabel('Absorbance') ## and here ofile = 'figures/examplespec.png' ## name of the output file fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resolution thing ### This part is for making this website, so you don't need it return ofile
simulated Beer's law data
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data C = [1.0, 2.0, 3.0, 4.0, 5.0] # these will be the x values in the graph they are simulating concentration in ppm A = [0.1, 0.2, 0.3, 0.4, 0.5] #this is the simulated Aborbance values ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ax.scatter(C,A, s=55,marker='.') ## ax is the variable for the axis where you want to work ## s is the size of the marker ## marker is the marker some choices: ('.','o', 'v', '^', '<', '>', '8', 's', 'p', '*', 'h', 'H', 'D', 'd', 'P', 'X') ax.set_xlabel('Concentration /ppm') ## latex works here ax.set_ylabel('Absorbance') ## and here ofile = 'figures/examplescatter.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Simulated Mass Data
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data mass = 10+np.random.randn(200)*0.29 ## these will be 200 simulated mass measurements normally distributed around 10 g ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ax.hist(mass) ## ax is the variable for the axis where you want to work ax.set_xlabel('Mass /g') ## latex works here ax.set_ylabel('Count') ofile = 'figures/examplehist.png' ## the output filename fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resolution thing ### This part is for making this website, so you don't need it return ofile
First some data
This part of computer programming is arguably the most difficult to learn for new coders. Some people call it data wrangling. These words mean getting the data from your instrument's output into the memory of python for processing. Because of this difficulty, in this course we will not deal with it. If you need help getting your data from your research into python for processing ask me, and we can do it outside of class. I am happy to assist you. I am making this teaching choice, because every instrument has its own kind of data file format, and to cover them all would take the entire semester.
So for this class, I will issue to you data in 3 kinds of formats:
- One will be single datafiles *.csv or *.dat with or without headers stored as text files. Some examples of types of data:
- chromatograms
- spectra
- concentration data
- A relative of the *.csv file is the MS Excel files.
It is pretty common for data/results to be stored in an Excel file. I believe this is because it is pretty easy to type data into spreadsheet programs such as MS Excel. I also believe that is why they are so popular. So we will learn how to read data from these file.
- The other type will be datasets in either *.mat files or *.hdf5 files. Some examples:
- hyphenated technique data (GC-MS, Raman spectroscopic imaging, 2D NMR data)
- lots of similar data (100's of IR spectra)
- This type of file is a kind of archive in which the data has been organized for you
- we will see some examples below
I will assume that all data is stored in a directory under your working directory named data. So all data file paths will look like
filename = 'data/somedatafilename.dat' # on mac or unix filename = 'data\fomedatafilename.dat' # on a windows computer
loading *.csv or *.dat files with no header and plot it
These will be files that contain 2 or more columns of data separated by a comma or tab or space. Traditionally, the first column is the independent values (you plot them on the x axis), and the 2nd, 3rd, etc. column(s) is(are) the dependent values (you plot them on the y-axis).
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data ifile = 'data/sample1.csv' ## the name and relative path of the data file rawdata = np.genfromtxt(ifile,delimiter=",") ## read the file output into an array print('the shape of the d is ',rawdata.shape) ## inspect the data size x = rawdata[:,0] ## extract/slice the wavenumbers (x values) spec = rawdata[:,1] ## extract/slice the absorbance values (y-values) ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ax.plot(x,spec) ## ax is the variable for the axis where you want to work ax.set_xlim((650,4000)) ## set graph's limits ax.invert_xaxis() ## vibrational spectra need this ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ## latex works here fig.tight_layout() ## removes a bunch of white space around figure ofile = 'figures/dataspec.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Loading *.mat files
This file format is matlab's high performance data format. It is based upon the HDF5 format. http://www.hdfgroup.org/HDF5/
This is the faster and easier of the file formats whenever you have lots of data to read into your program. The data is already stored in the variable. To use them, load this data.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module import scipy.io as sio #open matlab files module ## get data ifile = 'data/tolpxyl.mat' ## the path and name of the datafile mat_contents = sio.loadmat(ifile) ## read the file output into a dictionary print('The contents of the mat file is ',mat_contents.keys()) ##print the names of the contents of the matfile Aorg = mat_contents['Aorg'] ## retrieve the Spectra from the dictionary xorg = mat_contents['xorg'] ## retrieve the x values from the dictionary print('the shape of the Aorg is ',Aorg.shape) ## inspect the variable's size x = xorg.T ## xorg is stored funny in the mat file so we transpose it A = Aorg print('x and A shapes are ',x.shape,A.shape) ## process data
The contents of the mat file is dict_keys(['__header__', '__version__', '__globals__', 'Aorg', 'Ctrue', 'comp', 'xorg']) the shape of the Aorg is (9, 601) x and A shapes are (601, 1) (9, 601)
Reading data out of Excel files
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module import scipy.io as sio #open matlab files module ## get data ifile = 'data/glass_elemental_conc.xlsx' ## the path and name of the datafile df = pd.read_excel(ifile,index_col=0) print('df \n',df) ## process data
df RI Na Mg Al Si K Ca Ba Fe type SID 1 1.52101 13.64 4.49 1.10 71.78 0.06 8.75 0.00 0.0 1 2 1.51761 13.89 3.60 1.36 72.73 0.48 7.83 0.00 0.0 1 3 1.51618 13.53 3.55 1.54 72.99 0.39 7.78 0.00 0.0 1 4 1.51766 13.21 3.69 1.29 72.61 0.57 8.22 0.00 0.0 1 5 1.51742 13.27 3.62 1.24 73.08 0.55 8.07 0.00 0.0 1 .. ... ... ... ... ... ... ... ... ... ... 210 1.51623 14.14 0.00 2.88 72.61 0.08 9.18 1.06 0.0 7 211 1.51685 14.92 0.00 1.99 73.06 0.00 8.40 1.59 0.0 7 212 1.52065 14.36 0.00 2.02 73.42 0.00 8.44 1.64 0.0 7 213 1.51651 14.38 0.00 1.94 73.61 0.00 8.48 1.57 0.0 7 214 1.51711 14.23 0.00 2.08 73.36 0.00 8.62 1.67 0.0 7 [214 rows x 10 columns]
Loading Multiple Spectra and their Graphs
This idea means that you have to load more than one spectrum or chromatogram into memory for plotting. In this section we will first learn several techniques of loading multiple spectra, and then later we will plot them in several techniques. There are 3 basic strategies that can be employed to load and graph multiple spectra.
- manually if you just have a few spectra
- with a loop if you have lots of spectra or
- just load a *.mat file that I or someone else prepared for you.
The strategy that is used is somewhat subjective; meaning that we get to make the choice. I usually choose based upon my level of laziness at the moment.
For the demonstration of the different data loading strategies we will employee the spectral overlay graphing technique without any discussion of it. We will discuss the graphing a little later.
Multiple spectra: spectral overlays with just a few files loaded manually
There are two ways to do this. One is in x, y pairs.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data ## one datafile ifileA = 'data/sample1.csv' rawdataA = np.genfromtxt(ifileA,delimiter=",") ## read the file output into an array print('the shape of the d is ',rawdataA.shape) xA = rawdataA[:,0] ## extract the wavenumbers specA = rawdataA[:,1] ## extract the absorbance values ## second datafile ifileB = 'data/sample2.csv' rawdataB = np.genfromtxt(ifileB,delimiter=",") ## read the file output into an array print('the shape of the d is ',rawdataB.shape) xB = rawdataB[:,0] ## extract the wavenumbers specB = rawdataB[:,1] ## extract the absorbance values ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(xA,specA,label='spec 1') ## ax is the variable for the axis where you want to work ax.plot(xB,specB,label='spec 2') # do it again to plot the second spectrum ax.set_xlim((650,4000)) ## set graph's limits ax.invert_xaxis() ## vibrational spectra need this ax.legend() ## call the legond for the axis to show it ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ## latex works here ofile = 'figures/dataspecoverlay2.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Multiple spectra: loaded with a for loop (Just load and plot the data)
Notice in the code below that the data is not stored, but simply plotted. This might be useful, but does not allow the user to do anything else with all of these spectra.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## define some custom functions def dataloader(ifile): rawdata = np.genfromtxt(ifile,delimiter=",") ## read the file output into an array x = rawdata[:,0] ## extract/slice the wavenumbers (x values) spec = rawdata[:,1] return x,spec ## get data file list filelist = ['specfile1.csv', 'specfile2.csv', 'specfile3.csv', 'specfile4.csv', 'specfile5.csv'] ## list of files datahole = 'data/' ## load and plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure for ifile in filelist: ## loop through filelist name = ifile[:-4] ## create a legion name from file list x,spec = dataloader(datahole+ifile) ## use you new function to load the data;notice the concatantion of filename and path ax.plot(x,spec,label=name) ## plot each spectrum; add the label ax.set_xlim((650,4000)) ## set graph's limits ax.invert_xaxis() ## vibrational spectra need this ax.legend() ## call the legond for the axis to show it ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ## latex works here ofile = 'figures/dataspecoverlayforloop.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Multiple spectra: loaded with a for loop (Load, store and plot the data)
There are two ways to load the data, store it in an array-like object and plot. One uses an array, and the other uses either a dictionary or dataframe.
for loop: array data storage
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## define some custom functions def dataloader(ifile): rawdata = np.genfromtxt(ifile,delimiter=",") ## read the file output into an array x = rawdata[:,0] ## extract/slice the wavenumbers (x values) spec = rawdata[:,1] return x,spec ## get data file list filelist = ['specfile1.csv', 'specfile2.csv', 'specfile3.csv', 'specfile4.csv', 'specfile5.csv'] ## list of files datahole = 'data/' ## load and plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure A = np.zeros((5,3351)) ## pre populate the array with zeros for n,ifile in enumerate(filelist): ## loop through filelist; enumerate to create a counter name = ifile[:-4] ## create a legion name from file list x,spec = dataloader(datahole+ifile) ## use you new function to load the data;notice the concatantion of filename and path ax.plot(x,spec,label=name) ## plot each spectrum; add the label A[n,:] = spec ## store each spectrum in an array as they are loaded in the for loop ax.set_xlim((650,4000)) ## set graph's limits ax.invert_xaxis() ## vibrational spectra need this ax.legend() ## call the legond for the axis to show it ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ## latex works here ofile = 'figures/dataspecoverlayforlooparray.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
for loop: dictionary / dataframe data storage load file
In this technique the loading data and plotting are separated (2 for loops). They could be done together (1 for loop). This (2 for loop method) was chosen because it is easier to understand what is happening.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## define some custom functions def dataloader(ifile): rawdata = np.genfromtxt(ifile,delimiter=",") ## read the file output into an array x = rawdata[:,0] ## extract/slice the wavenumbers (x values) spec = rawdata[:,1] return x,spec ## multiple files from an excel file datahole = 'data/' filelist = ['specfile1.csv', 'specfile2.csv', 'specfile3.csv', 'specfile4.csv', 'specfile5.csv'] ## load the data datadict = {} ## create an empty dictionary to store the data for ifile in filelist: ## for loop through the filelist x,spec = dataloader(datahole+ifile) key = ifile[:-4] ## get dictionary key from ifile datadict[key] = spec ## store spectrum in dictiorary df = pd.DataFrame.from_dict(datadict,orient='index',columns=x) ## convert dictionary to dataframe fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure #for key in df.index: ## loop through dataframe one index at a time # x = df.columns ## extract the xvalues from the columns of the dataframe # sp = df[df.index == key].values.T ## get the spectrum that correponds to the indexed spectrum (key) # ax.plot(x,sp,label=key) ##plot that spectrum df.T.plot(ax=ax) ## plot the dataframe with index as labels on the axis ax ax.set_xlim((650,4000)) ## set graph's limits ax.invert_xaxis() ## vibrational spectra need this ax.legend() ## call the legond for the axis to show it ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ## latex works here ofile = 'figures/dataspecoverlayforloopdict.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
for loop: dictionary / dataframe data storage load file information from Excel
In this technique the loading data and plotting are separated (2 for loops). They could be done together (1 for loop). This (2 for loop method) was chosen because it is easier to understand what is happening. We will also get the file information from an MS Excel file. This is very sophisticated and we will not rely on this for our homework.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## define some custom functions def dataloader(ifile): rawdata = np.genfromtxt(ifile,delimiter=",") ## read the file output into an array x = rawdata[:,0] ## extract/slice the wavenumbers (x values) spec = rawdata[:,1] return x,spec ## multiple files from an excel file datahole = 'data/' mdf = pd.read_excel('data/metadatafile.xlsx',index_col=0) ## read file info from Excel meta data file #mdf ## load the data datadict = {} ## create an empty dictionary to store the data for ifile in mdf.index: ## for loop through the mdf.index x,spec = dataloader(datahole+ifile) ## load data key = ifile[:-4] ## extract key and legend name from the mdf datadict[ifile] = spec ## store the spectrum in the dictionary under the key # the indices of the dataframe are the keys in the dicitonary (or part of the filenames) # the columns of the df are the wavenumbers of the spectra ## all spectra must be the same size for this to work df = pd.DataFrame.from_dict(datadict,orient='index',columns=x) ## convert the dictionary to a dataframe # plot the data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure #for ifile in df.index: ## for loop through each index # x = df.columns ## if you already have the x-values in memory then you don't need to do this step # sp = df[df.index == ifile].values.T # get the spectra out one index at a time # name = mdf[mdf.index == ifile].values[0][0] ## get the legend name from mdf # ax.plot(x,sp,label=name) ## plot each spectrum labeled with its name leglist = mdf.key.values.tolist() ## extract the legend labels from the mdf fig,ax = plt.subplots(figsize= (12,4)) ## create an axis and figure df.T.plot(ax=ax) ## plot the dataframe with index as labels ax.legend(leglist) ## call the legond for the axis to show it ax.set_xlim((650,4000)) ## set graph's limits ax.invert_xaxis() ## vibrational spectra need this #ax.legend() ## call the legond for the axis to show it ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ## latex works here ofile = 'figures/dataspecoverlayforloopdict_fromexcel.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Multiple Spectra: from a pre-made dataset file (HDF or mat)
When you have an inconvenient number of spectra to load manually, a good choice is to load a dataset. Two good choices of dataset file formats are hdf5 files and matlab files (which are actually just hdf5 files).
In this case, these datasets need to be made either by the computer software attached to the instrument that acquired the data or separately in another program. In the latter case, this process is often a custom job, so we will not show it here, but involves looping through every spectrum, storing them in an array and then saving arrays as a file. All of the spectra must be the same size for this to work.
Often these type of data don't get a legend, because the number of spectra would make the legend size ridiculous. This number of spectra to plot is a pretty good delimiter for these two types of data.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module import scipy.io as sio #module for opening matlab files ## get data ifile = 'data/tolpxyl.mat' mat_contents = sio.loadmat(ifile) ## read the file output into a dictionary print(mat_contents.keys()) ##print the names of the contents of the matfile Aorg = mat_contents['Aorg'] ## retrieve the Spectra xorg = mat_contents['xorg'] ## retrieve the x values print('the shape of the Aorg is ',Aorg.shape) x = xorg.T A = Aorg ## process data ## plot data fig, ax = plt.subplots(figsize=(9,6)) ## figsize is the aspect ratio of the figure ax.plot(x,A.T) ## ax is the variable for the axis where you want to work ax.set_xlim((200,340)) ## set graph's limits #ax.invert_xaxis() ## vibrational spectra need this ax.set_xlabel('Wavelength /nm') ## latex works here ax.set_ylabel('Absorbance') ## latex works here ofile = 'figures/multidataspecoverlay.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
- ASS10 Assignment 10 loading some data (no plotting of data required)
NOTE: for this assignment all of the files will be in the ass10 folder not in the ass10/data folder to prevent conflict between Microsoft Windows and MacOS
- Part 1
Create a function that will
- load a single data file (xy columns) from a path + filename input
- return the x values and y values as 2 separate 1D arrays
- Assumed that the data files will be txt files of 2 columns of data each separated by comma (a CSV format)
- HINT: we did this in the videos
- Part 2
Create a function that will use the function in ass10 Part 1 to
- load a set of data files (use the function from part 1 here) from a list of file names as the input
- store the x-values into a 1D np.array
- store the y-values into an 2D np.array
- the number of rows should be the number of files that you load
- the number of columns should be the number of y-values per file
- return the x-values as a 1D array and the y-values a 2D array
- Part 1
Spectra plotting (overlayed)
This method works well if you have spectra that are either the same size as each other or different sizes (number of wavelengths). One limitation is that you overwrite the spectral data as you step through the loop. So as this code stands you can not process the data.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data def dataloader(ifile): import numpy as np rawdata = np.genfromtxt(ifile, delimiter=',') ## read the file output into an array x = rawdata[:,0] ## extract the wavenumbers spec = rawdata[:,1] ## extract the absorbance values return x, spec ## return the wavenumbers and absorbance values ifile1 = 'sample2.csv' ## datafile named stored in variable ifile2 = 'specfile1.csv' ## datafile named stored in variable datahole = 'data/' #location of datafiles (path) x,spec1 = dataloader(datahole + ifile1) ## load data file notice the concatenation of the path and datfile x,spec2 = dataloader(datahole + ifile2) fig, ax = plt.subplots(figsize=(9,6)) ## create a figure and axis to plot onto ax.plot(x,spec1,label='spec 1') ## plot one file with label ax.plot(x,spec2 ,label='spec 2') ## plot another file with label ax.set_xlim((700,2000)) ## change the view of the graph ax.invert_xaxis() ##important for IR ax.legend() ## show the legend ax.set_ylabel('Absorbance') ## set the y axis label ax.set_xlabel('Wavenumber /cm$^{-1}$') ## set the x axis label ax.grid() ## turn on a basic grid to help see which bands lineup # put a vertical line on the ax axis # linestyles [‘solid’ | ‘dashed’, ‘dashdot’, ‘dotted’ | (offset, on-off-dash-seq) | '-' | '--' | '-.' | ':' | 'None' | ' ' | ''] # alpha is transparency 0 - 1 ax.axvline(x=837,linestyle='--',color='g',alpha=0.5) ofile = 'figures/dataspecoverlay.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
spectra plotting (stacked)
Sometimes overlaying spectra can confuse the viewer of the data. So another way to display the data is to stack them.
Whenever you stack spectra, you need to keep somethings in mind. One, the y axis is no longer accurate, so you need to indicate this to the viewer. A customary way to do this is to add a vertical bar on the graph that indicates a specific δ y distance.
Since the data are not overlayed it is difficult for the viewer to determine if spectral features overlap. So to indicate the x position of a particular band or peak, you will need to annotate the graph. There are several ways to do this:
- add grid lines (this works well if you are printing the graph and there are lots of bands that you want the viewer to compare)
- add a band label with position and/or height
- add a vertical line centered on just a few of the bands
Look in the tricks below to find the details of all of these.
Here is a basic example of a stacked plot. A constant is added to one of the y vector to move it up.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data def dataloader(ifile): import numpy as np rawdata = np.genfromtxt(ifile, delimiter=',') ## read the file output into an array x = rawdata[:,0] ## extract the wavenumbers spec = rawdata[:,1] ## extract the absorbance values return x, spec ## return the wavenumbers and absorbance values ifile1 = 'sample2.csv' ## datafile named stored in variable ifile2 = 'specfile1.csv' ## datafile named stored in variable datahole = 'data/' #location of datafiles (path) x,spec1 = dataloader(datahole + ifile1) ## load data file notice the concatenation of the path and datfile x,spec2 = dataloader(datahole + ifile2) offset = 0.2 fig, ax = plt.subplots(figsize=(9,6)) ## create a figure and axis to plot onto ax.plot(x,spec1,label='spec 1') ## plot one file with label ax.plot(x,spec2 + offset ,label='spec 2') ## plot another file with label notice the offset ax.set_xlim((700,2000)) ## change the view of the graph ax.invert_xaxis() ##important for IR ax.legend() ## show the legend ax.set_ylabel('Absorbance') ## set the y axis label ax.set_xlabel('Wavenumber /cm$^{-1}$') ## set the x axis label ax.grid() ## turn on a basic grid to help see which bands lineup ax.set_yticklabels('') # absorbance is no longer correct so turn it off # add a bar that indicates the absorbance scale for the two spectra ax.annotate('0.1 absorbance',(1800,0.20),(1800,0.3),ha='center',arrowprops={'arrowstyle':'|-|'}) # put a vertical line on the ax axis # linestyles [‘solid’ | ‘dashed’, ‘dashdot’, ‘dotted’ | (offset, on-off-dash-seq) | '-' | '--' | '-.' | ':' | 'None' | ' ' | ''] # alpha is transparency 0 - 1 ax.axvline(x=837,linestyle='--',color='g',alpha=0.5) ofile = 'figures/dataspecstack.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
scatter plot of linear data
Whenever you have data with limited length (number of points < 50), it is probably better to graph it as a scatter plot than a line plot. If you have only one data set this is pretty easy.
Only one data set
Just hard code into the program the data.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data C = [1, 2, 3, 4, 5] # these will be the x values in the graph they are simulating concentration in ppm A = [0.1, 0.2, 0.3, 0.4, 0.5] #this is the simulated Aborbance values ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ax.scatter(C,A) ## ax is the variable for the axis where you want to work ax.set_xlabel('Concentration /ppm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplescatter.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Multiple sets of data
If the number of sets of data are relatively small, the you can also hard code them into the code.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data toluene and xylone C1 = [1, 2, 3, 4, 5] # these will be the x values in the graph they are simulating concentration in ppm A1 = [0.1, 0.2, 0.3, 0.4, 0.5] #this is the simulated Aborbance values C2 = [0.98, 2.10, 3.05, 4.09, 5.01] # toluene A2 = [0.11, 0.25, 0.39, 0.51, 0.691] # toluene ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ax.scatter(C1,A1,label='p-Xylene') ## ax is the variable for the axis where you want to work ax.scatter(C2,A2,label='toluene') ax.legend() ## turn the legend on ax.set_xlabel('Concentration of analyte in water /ppm') ## latex works here ax.set_ylabel('Absorbance at 252 nm') ofile = 'figures/examplescattermulti.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
histogram of data
Here are some final exam grades from a class of mine. Notice that the grades are also hard coded. This is a nice way to do it, because the data and the code travel together, which makes the execution of the code very reproducible.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data # here are some final grades of mine g = [59.87, 65.35, 53.56, 35.73, 46.15, 59.05, 43.27, 67.27, 50.27, 46.70, 43.41, 62.61, 50.27, 66.73, 61.79, 44.78, 43.41, 50.82, 48.07, 49.17, 49.72, 57.13, 44.51, 47.53, 49.17, 70.02, 61.24, 48.35, 18.73, 62.61, 40.94, 56.85, 43.14, 55.75, 62.61, 55.48, 46.15, 43.41, 54.11] ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ax.hist(g) ax.set_xlabel('Grades') ## latex works here ax.set_ylabel('Count') ofile = 'figures/examplehist.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Tricks that can be applied to most or all of the Former Graphs
Computer Screen Inspection of Data
This is what you get as the default.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data x = np.arange(200,400) # these will be the x values in the graph they are simulating a UV spectrum xc = 250 # this is the center of the band sigma = 10 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(x,spec) ## ax is the variable for the axis where you want to work ax.set_xlabel('Wavelength /nm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespec.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Publication Quality Printout of Data
This is basically controlled by the output method that you choose. In the fig.savefig command the keyword dpi stands for dots per inch. The bigger the dpi number the higher resolution. If you pick a raster graphic file (tiff,jpg,png) the bigger the number after "dpi=" the more pixels per inch in the final file. If you pick a vector graphics type (eps,svg, ps) then the resolution parameter is useless. Publication quality figures usually have a dpi of 600 to 1200. The publisher will tell you what they want.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data x = np.arange(200,400) # these will be the x values in the graph they are simulating a UV spectrum xc = 250 # this is the center of the band sigma = 10 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(x,spec) ## ax is the variable for the axis where you want to work ax.set_xlabel('Wavelength /nm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespecpub.png' fig.savefig(ofile,dpi=300) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Thumbnail or Lab Notebook Printout of Data
Since this figure does not need to be as high quality, you can simply change the resolution parameter dpi from 300 to 75.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data x = np.arange(200,400) # these will be the x values in the graph they are simulating a UV spectrum xc = 250 # this is the center of the band sigma = 10 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(x,spec) ## ax is the variable for the axis where you want to work ax.set_xlabel('Wavelength /nm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespecthumb.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Beamer Presentation of Data
When you are projecting data for a conference or seminar the projectors are the limiting factor in data display. Since this figure does not need to be as high quality, you can simply change the resolution parameter dpi from 300 to 150-100.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data x = np.arange(200,400) # these will be the x values in the graph they are simulating a UV spectrum xc = 250 # this is the center of the band sigma = 10 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(x,spec) ## ax is the variable for the axis where you want to work ax.set_xlabel('Wavelength /nm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespecbeamer.png' fig.savefig(ofile,dpi=150) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Reversing the x-axis for Vibrational Spectroscopic Data
According the IUPAC standards (PAC, 2007 pg 354) for spectra, higher energy should be plotted on the left side of a spectra. For vibrational spectra with wavenumber axis tick labels running between 400 and 4000 cm-1 (low to high), traditional graphics programs plot the spectra backwards. Python provides a simple means of reversing the x-axis direction (or y-axis if need be). Here is an example:
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data x = np.arange(1600,1800) # these will be the x values in the graph they are simulating a UV spectrum xc = 1656 # this is the center of the band sigma = 50 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(x,spec) ## ax is the variable for the axis where you want to work ax.invert_xaxis() ## vibrational spectra need this ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespecreversex.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Change the Line Width of the Lines so the Show Up Better
On the computer screen using python, it is relatively easy to see the lines of a graph. When you display them in a beamer presention or even in publication quality graphics, it is not always so clear. You can control the linewidth.
Here is an example controlling the linewidth in the plot command:
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module lw = 5 # use a variable to control graphs at once ## get data x = np.arange(1600,1800) # these will be the x values in the graph they are simulating a UV spectrum xc = 1656 # this is the center of the band sigma = 50 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ax.plot(x,spec, linewidth=lw) ## ax is the variable for the axis where you want to work ax.invert_xaxis() ## vibrational spectra need this ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespeclinewidth.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Here is an example controlling the linewidth for the whole Python or Jupyter session:
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ###these next lines write the parameters lw,lsz to plotting parameters import matplotlib as mpl ## this loads the moduule matplotlib as mpl lw = 5 #this is the linewidth lsz = 12 #this is the font size mpl.rcParams['xtick.labelsize'] = lsz mpl.rcParams['ytick.labelsize'] = lsz mpl.rcParams['font.size'] = lsz mpl.rcParams['lines.linewidth'] = lw ## get data x = np.arange(1600,1800) # these will be the x values in the graph they are simulating a UV spectrum xc = 1656 # this is the center of the band sigma = 50 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(5,4)) ## figsize is the aspect ratio of the figure ax.plot(x,spec) ## ax is the variable for the axis where you want to work ax.invert_xaxis() ## vibrational spectra need this ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespeclinewidth2.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Change or Control the Color of the Line
You can control the color of a line in a similar method as you control the linewidth.
| RGBValue | Short Name | Long Name |
|---|---|---|
| (1, 1, 0) | y | yellow |
| (1, 0, 1) | m | magenta |
| (0, 1, 1) | c | cyan |
| (1, 0, 0) | r | red |
| (0, 1, 0) | g | green |
| (0, 0, 1) | b | blue |
| (1, 1, 1) | w | white |
| (0, 0, 0) | k | black |
Here is a basic or simple example of controlling the color of a line:
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data x = np.arange(200,400) # these will be the x values in the graph they are simulating a UV spectrum xc = 350 # this is the center of the band sigma = 50 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(6,5)) ## figsize is the aspect ratio of the figure ax.plot(x,spec,'r') ## ax is the variable for the axis where you want to work ax.plot(x,spec+0.5,'k') ax.set_xlabel('Wavelength /nm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespeccolor1.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
For finer color control, use a color tuple (3 number sequence surrounded by ( ) to get definition of the color.
Usage: (R, G, B) the numbers are between 0 and 1. All ones (1,1,1) is white, all zeros is black (0,0,0), and (1,0,0) is red.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module # use a color tuple # this is a dictionary of colors usage cl[0] yields blue (0,0,1) # you don't have to use this dictionary, I just like it. # you could also redefine the keys to something more memorable cl = {0:(0, 0, 1),#; %blue 4:(0, 1, 0),#; % bright green 6:(0, 1, 1),#; % cyan 3:(0, 0, 0),#; %black 1:(1, 0, 0),#; % red 7:(1, 1, 0),# % yellow 2:(1, 0, 1),#; % magenta 5:(1, 0.5, 0.5),#; % pink brown 8:(0.75, 1, 0.75),# % dull light green 9:(0.5, 0.5, 1),# % greyblue 10:(1, 0.5, 0),#; % orange 11:(0, 0.8, 0.5),# % bluegreen 12:(0.5, 0, 0.6),# % purple 13:(1,0.75,0),## orange 14:(0.5,0,1), ##also purple, dark 15:(0.78,1,0), ##ugly green 16:(0,0.5,1), #blue ski 17:(1,0,0.25), #pink red 18:(0.39,0.39,0.49), #grey blue 19:(0.02,0.42,0.1), #forest green 20:(0.75,0.26,0.36) #dark red } ## get data x = np.arange(200,400) # these will be the x values in the graph they are simulating a UV spectrum xc = 350 # this is the center of the band sigma = 50 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data color1 = cl[0] ## get the 0th color tuple color2 = cl[1] ## get the 1st color tuple print(color1, color2) fig, ax = plt.subplots(figsize=(6,4)) ## figsize is the aspect ratio of the figure ax.plot(x,spec,color=color1) ## ax is the variable for the axis where you want to work ax.plot(x,spec+0.5,color=color2) ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplespeccolor1.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
In the past there was a lot of expense associated with using color in journal/report figures, so many people avoided using color. Today these ideals are much less important. However, color should not be used just for the sake of using color. Your color choices should enhance understanding of the graphics, not just prettyfy them.
Colors and markers are attributes that can be controlled in figures. Here is a table of markers.
| specifier | Marker Type |
|---|---|
| '+' | Plus sign |
| 'o' | Circle |
| '*' | Asterisk |
| '.' | Point |
| 'x' | Cross |
| 'square' or 's' | Square |
| 'diamond' or 'd' | Diamond |
| '^' | Upward-pointing triangle |
| 'v' | Downward-pointing triangle |
| '>' | Right-pointing triangle |
| '<' | Left-pointing triangle |
| 'pentagram' or 'p' | Five-pointed star (pentagram) |
| 'hexagram' or 'h' | Six-pointed star (hexagram) |
Here is an example of controlling both the color and the marker in a scatter plot.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data C = [1, 2, 3, 4, 5] # these will be the x values in the graph they are simulating concentration in ppm A = np.array([0.1, 0.2, 0.3, 0.4, 0.5]) #this is the simulated Aborbance values ## process data ## plot data fig, ax = plt.subplots(figsize=(6,5)) ## figsize is the aspect ratio of the figure ax.scatter(C,A,marker='o',color='r',label='data 1') ## ax is the variable for the axis where you want to work ax.scatter(C, A+0.5,marker='+',color='r',label='data 2') ## ax is the variable for the axis where you want to work ax.scatter(C, A+0.3,marker='x',color='g',label='data 3') ax.legend() ax.set_xlabel('Concentration /ppm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplescattermakers.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resolution thing ### This part is for making this website, so you don't need it return ofile
Adding Axis Labels
Use the axis methods .setxlabel('your words here') and .setylabel('your label here') to add/change the labels on axes.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data C = [1, 2, 3, 4, 5] # these will be the x values in the graph they are simulating concentration in ppm A = np.array([0.1, 0.2, 0.3, 0.4, 0.5]) #this is the simulated Aborbance values ## process data ## plot data fig, ax = plt.subplots(figsize=(6,4)) ## figsize is the aspect ratio of the figure ax.scatter(C,A,marker='o',color='r') ## ax is the variable for the axis where you want to work ax.set_xlabel('Concentration /M') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/exampleaxislabels.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
control the font size
To control the font:
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ###these next lines write the parameters lw,lsz to plotting parameters for the whole jupyter/python session import matplotlib as mpl ## this loads the module matplotlib as mpl lw = 5 #this is the linewidth lsz = 50 #this is the font size, 50 is probably WAAAAAYYYYY too big. mpl.rcParams['xtick.labelsize'] = lsz mpl.rcParams['ytick.labelsize'] = lsz mpl.rcParams['font.size'] = lsz mpl.rcParams['lines.linewidth'] = lw ## get data x = np.arange(1600,1800) # these will be the x values in the graph they are simulating a UV spectrum xc = 1656 # this is the center of the band sigma = 50 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(x,spec) ## ax is the variable for the axis where you want to work ax.invert_xaxis() ## vibrational spectra need this ax.set_xlabel('Wavenumber /cm$^{-1}$') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplesfont1.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
insert a chemical structure into the graphic
- State "TODO" from
Not done yet.
Multiple Graphs on one Figure
To add multiple graphs on one figures use the subplot command to get a different axes for each graph.
subplot(numberrows,numbercolumns,plotindex)
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data x = np.arange(200,400) # these will be the x values in the graph they are simulating a UV spectrum xc = 250 # this is the center of the band sigma = 30 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig = plt.figure(figsize=(12,8)) ax1 = fig.add_subplot(211) ## axis 1 for subplot 1 ax1.plot(x,spec) ax1.set_xlabel('Wavelength /nm') ## latex works here ax1.set_ylabel('Absorbance') ax2 = fig.add_subplot(212) ## axis 2 for subplot 2 ax2.plot(x,np.sin(x)) ax2.set_xlabel('Time /min') ax2.set_ylabel('Aplitude /V') ofile = 'figures/examplesuplots.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Resize the Displayed Region (zoom)
Sometimes the information that you want to display is not easy to see if you show all of your data at once. In these situations it is better zoom in on the regions of interest. If you want to show where the band max is for example.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module from matplotlib.patches import Rectangle ## for drawing the rectangle ## get data x = np.arange(200,400,0.1) # these will be the x values in the graph they are simulating a UV spectrum xc = 250 # this is the center of the band sigma = 30 # %sigma*2*sqrt(2*ln(2)) is the width at half height of the Y max of the curve spec = np.exp(-1*((x-xc)**2)/(2*sigma**2)) ##this will make a Gaussian curve ## process data ## plot data fig = plt.figure(figsize=(12,8)) ax0 = fig.add_subplot(211) ## whole specturm ax0.plot(x,np.sin(x)) ax0.set_xlabel('Wavelength /nm') ## latex works here ax0.set_ylabel('Absorbance') ax0.add_patch(Rectangle((225,0.4),width=50, height=0.6,alpha=0.5)) ax1 = fig.add_subplot(212) ax1.plot(x,np.sin(x)) ax1.set_xlabel('Wavelength /nm') ## latex works here ax1.set_ylabel('Absorbance') ax1.set_xlim(225,275) ### control the view of lower graph ax1.set_ylim(0.4, 1) ### control the view ofile = 'figures/examplezoom.png' fig.savefig(ofile,dpi=75) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
Adding a Vertical Line to a Graph and Text annotations
If you want to indicate where a band is in more than one spectrum, a nice way to do this is to draw a vertical line though the graph.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module import scipy.io as sio #open matlab files ## get data ifile = 'data/tolpxyl.mat' mat_contents = sio.loadmat(ifile) ## read the file output into a dictionary print(mat_contents.keys()) ##print the names of the contents of the matfile Aorg = mat_contents['Aorg'] ## retrieve the Spectra xorg = mat_contents['xorg'] ## retrieve the x values print('the shape of the Aorg is ',Aorg.shape) x = xorg.T A = Aorg C = mat_contents['Ctrue'] ## concentration matrix ## process data ## plot data fig = plt.figure(figsize=(12,8)) ## figsize is the aspect ratio of the figure ax1 = fig.add_subplot(211) ## ax1 is the variable for the axis where you want to work ax1.plot(x,A.T) ax1.set_xlim((200,340)) ## set graph's limits # put a vertical line on the ax axis # linestyles [‘solid’ | ‘dashed’, ‘dashdot’, ‘dotted’ | (offset, on-off-dash-seq) | '-' | '--' | '-.' | ':' | 'None' | ' ' | ''] # alpha is transparency 0 - 1 ax1.axvline(x=222.5,linestyle='--',color='k',alpha=0.9) ax1.axvline(x=274,linestyle='--',color='r',alpha=0.9) ax1.text(224, 2.62, ' this band is not linear with concentration') ## text annotations ax1.text(280,1.3, '$\leftarrow$ this band is linear with concentration') ##latex annotations ax1.set_xlabel('Wavelength /nm') ## latex works here ax1.set_ylabel('Absorbance') ax2 = fig.add_subplot(212) ax2.scatter(C[:,0],A[:,304],color='r',label='274 nm') ax2.scatter(C[:,0],A[:,509],color='k',label='222 nm') ax2.set_xlabel('Analyte Concentration v/v /%') ## latex works here ax2.set_ylabel('Absorbance') ax2.legend() ofile = 'figures/exampleverticalline.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
adding grid lines
Grid lines also allow your reader to interpret the data in a graph. They are particularly useful for ovelaping data sets.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module import scipy.io as sio #open matlab files ## get data ifile = 'data/tolpxyl.mat' mat_contents = sio.loadmat(ifile) ## read the file output into a dictionary print(mat_contents.keys()) ##print the names of the contents of the matfile Aorg = mat_contents['Aorg'] ## retrieve the Spectra xorg = mat_contents['xorg'] ## retrieve the x values print('the shape of the Aorg is ',Aorg.shape) x = xorg.T A = Aorg ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(x,A.T) ## ax is the variable for the axis where you want to work ax.set_xlim((200,340)) ## set graph's limits # which ---> major, minor, both # Turn on the minor TICKS, which are required for the minor GRID ax.minorticks_on() ax.grid(True, which='major') ##select which to show, 'major' grid lines or 'minor' ax.set_xlabel('Wavelength /nm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/examplegrid.png' fig.savefig(ofile,dpi=300) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
adding a band label
Another way of indicating areas of the graph that you want to talk about in the text or bring into focus for the user is to add textual annotations such as a label to your graph. In python you can do this by importing your graph into powerpoint and manually add the label. This works fine until your professor makes you recreate every figure using the Greek alphabet. If you are finding that you need to recreate lots of figures and/or don't like using a mouse you can also add annotations programatically. Here is an example.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module import scipy.io as sio #open matlab files ## get data ifile = 'data/tolpxyl.mat' mat_contents = sio.loadmat(ifile) ## read the file output into a dictionary print(mat_contents.keys()) ##print the names of the contents of the matfile Aorg = mat_contents['Aorg'] ## retrieve the Spectra xorg = mat_contents['xorg'] ## retrieve the x values print('the shape of the Aorg is ',Aorg.shape) x = xorg.T A = Aorg ## process data ## plot data fig, ax = plt.subplots(figsize=(12,9)) ## figsize is the aspect ratio of the figure ax.plot(x,A.T) ## ax is the variable for the axis where you want to work ax.set_xlim((200,340)) ## set graph's limits ax.text(275,1.2, '(a)',fontsize=25) ## this fontsize is too big ax.set_xlabel('Wavelength /nm') ## latex works here ax.set_ylabel('Absorbance') ofile = 'figures/exampletext.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
adding a scale bar to the graph
I don't really have a great way of doing this. I will think about it and add it later.
# many functions are avaible in modules or libraries # in this example we will load the numpy module of functions import numpy as np #this command loads all of the functions in numpy and labels them np import pandas as pd # data organization module import matplotlib.pyplot as plt # plotting module ## get data def dataloader(ifile): import numpy as np rawdata = np.genfromtxt(ifile, delimiter=',') ## read the file output into an array x = rawdata[:,0] ## extract the wavenumbers spec = rawdata[:,1] ## extract the absorbance values return x, spec ## return the wavenumbers and absorbance values ifile1 = 'sample2.csv' ## datafile named stored in variable datahole = 'data/' #location of datafiles (path) x,spec1 = dataloader(datahole + ifile1) ## load data file notice the concatenation of the path and datfile fig, ax = plt.subplots(figsize=(12,9)) ## create a figure and axis to plot onto ax.plot(x,spec1,label='spec 1') ## plot one file with label ax.set_xlim((700,2000)) ## change the view of the graph ax.invert_xaxis() ##important for IR ax.legend() ## show the legend ax.set_ylabel('Absorbance') ## set the y axis label ax.set_xlabel('Wavenumber /cm$^{-1}$') ## set the x axis label ax.annotate('0.05 abosorbance',(1750,0.0),(1750,0.05), arrowprops={'arrowstyle':'|-|'},ha='center') ofile = 'figures/dataspeccalebar.png' fig.savefig(ofile,dpi=100) ## dpi is dots per inch and is a figure resotution thing ### This part is for making this website, so you don't need it return ofile
ass11
- create a Jupyter notebook (use the blank ass11.ipynb already in the directory) from scratch in the ass11 directory
- load one or more data file(s) from your ass11/data directory
- plot the data properly
- use 2 of the following tricks
- make a stacked plot
- added grid lines
- use subplots to create 2 axes
- add a bandlabel
- add a vertical line to highlight one band
- full credit will only be given to graphs (located in the figures directory) that are perfect, partial credit will be given.
- No credit without showing your work (code) in the ass11.ipynb
- make a 150 dpi graph of the data
- save this file into the ass11/figures directory