An unprecedented wealth of data is being generated by large genome, metagenome, and epigenetic projects, as well as other efforts to determine the structure and function of molecular biological systems. This technical elective focuses on algorithms and data structures for analyzing biomolecular data. In other words, CS 144 is a data science course oriented toward biomolecular data.
Catalog description
Introduces fundamental algorithms and data structures for solving analytical problems in molecular biology and genomics, including exact and approximate string matching, sequence alignment, genome assembly, and recognition of genes and regulatory motifs.
Credit is awarded for one of CS 144, CS 234, or CS 238.
Prerequisites
CS 141
Solid programming experience, ideally in Python
Basic probability and statistics
No biology background is assumed
Course information
Lectures
Tuesday & Thursday, 9:30–10:50 a.m., Boyce Hall 1471. Note: Lectures will not be recorded.
Discussions
Wednesdays, 3:00–3:50 p.m. over Zoom. Meeting details will be posted on Canvas.
Discussion forum
We will use a Discord server for discussion and questions about CS 144. The instructor will moderate the forum and respond to questions, and students are encouraged to help one another through discussion. Do not discuss assignment-specific solutions. See Canvas for the invitation, and please be respectful.
Office hours
Fridays, 3:00–3:50 p.m. over Zoom. Meeting details will be posted on Canvas. Meetings are also available by appointment.
Introduction to molecular and computational biology, including biotech tools
Space-efficient data structures for sequences
Short-read mapping: suffix tries and trees, suffix arrays, and the Burrows–Wheeler transform
Global, local, linear-space, and multiple sequence alignment
Genome assembly, overlap graphs, and de Bruijn graphs
Hidden Markov models, profile HMMs, Viterbi, and Baum–Welch learning
Motif finding and Gibbs sampling
Construction of evolutionary trees (phylogeny)
Coursework and policies
Course format
Seven individual homework assignments developed in JupyterLab: 50% of the grade
One programming project: 50% of the grade
Academic integrity
Cheating will not be tolerated. Homework and the final project must be completed independently without relying on AI tools. You may use only the external sources listed on this page or explicitly allowed by the instructor. Do not submit answers or code that you did not write yourself. Violations will receive a zero on the assignment and may receive a zero in the course, depending on severity, and will be referred to Student Conduct and Academic Integrity Programs. If you are unsure whether something is allowed, ask before submitting.
Late work
Each student receives five late days, usable in whole-day increments on any homework assignment. For more serious circumstances, contact the instructor.
Homework will be released as Python notebooks on Sundays in the Assignments area of Canvas and will be due the following Sunday at 11:59 p.m. Download each notebook, upload it to the department's JupyterLab server, and complete your work there. Submit the finished notebook on Canvas by the deadline. Solutions will be posted on Canvas.
Fall 2026
Course calendar
Schedule subject to change; updates will be announced on Canvas.
Opening day
Introduction and molecular biology
Week 1
Molecular biology
Molecular biology
Homework 1 posted
Week 2
Molecular biology
Read mapping
Homework 1 due · Homework 2 posted
Week 3
Read mapping
Read mapping
Homework 2 due · Homework 3 posted
Week 4
Programming project discussion
Sequence alignment
Homework 3 due · Homework 4 posted
Week 5
Sequence alignment
Genome assembly
Homework 4 due
Week 6
Genome assembly
Hidden Markov models
Homework 5 posted
Week 7
Hidden Markov models
Hidden Markov models
Homework 5 due · Homework 6 posted
Week 8
Motif finding
Motif finding
Homework 6 due · Homework 7 posted
Week 9
Evolutionary trees
No class · Thanksgiving holiday
Homework 7 due
Week 10
Evolutionary trees
Evolutionary trees and course wrap-up
Programming project due
Finals week
Project demonstrations
Additional references
Richard Durbin, A. Krogh, G. Mitchison, and S. Eddy, Biological Sequence Analysis: Probabilistic Models of Proteins and Nucleic Acids, Cambridge University Press, 1999.
Dan Gusfield, Algorithms on Strings, Trees and Sequences: Computer Science and Computational Biology, Cambridge University Press, 1997.
Dan E. Krane and Michael L. Raymer, Fundamental Concepts of Bioinformatics, Benjamin Cummings, 2002.
Neil C. Jones and Pavel Pevzner, An Introduction to Bioinformatics Algorithms, MIT Press, 2004.
Marketa Zvelebil and Jeremy O. Baum, Understanding Bioinformatics, Garland Science, 2007.
Additional resources
Foldit, a protein-folding game that contributes to scientific research