Week 2 Lecture prep
Homework Questions from Professor Jacobson
1. Nature’s machinery for copying DNA is called polymerase. What is the error rate of polymerase? How does this compare to the length of the human genome. How does biology deal with that discrepancy?
According to Albertson & Preston (2006), it is estimated that replicative DNA polymerases make errors approximately once every 10⁴–10⁵ nucleotides polymerized; i.e. the error rate is 10⁻⁴ to 10⁻⁵, before proofreading and post‑replicative repair. The human genome is roughly 3 billion base pairs (3 × 10⁹) (Cooper, 2000, The Human Genome).
If DNA polymerase worked alone with an error rate of 10⁻⁴ to 10⁻⁵, each time a cell divided it would introduce:
- At 10⁻⁵ error rate: ~30,000 errors per genome replication
- At 10⁻⁴ error rate: ~300,000 errors per genome replication
How biology solves this: cells use a three-tier system that works sequentially:
- Nucleotide selectivity (~10⁻⁵ error rate): The polymerase active site favors correct base pairing through shape complementarity and hydrogen bonding (reference)
- Exonuclease proofreading (~100-1000× improvement): The 3’→5’ exonuclease activity catches errors immediately after incorporation, removing mismatched nucleotides before continuing (reference)
- Mismatch repair (~another 100-1000× improvement)
Together, these mechanisms reduce the effective error rate to roughly 10⁻⁹ – 10⁻¹⁰ per base (i.e 0.3 - 3 mutations per genome per replication), so only a few mutations are fixed per cell division despite the genome’s size. (Kunkel, 2009)
References:
Albertson, T. M. & Preston, B. D., 2006. DNA replication fidelity: proofreading in Trans. Current Biology, 16(6), pp.R209–R211. Available at: https://doi.org/10.1016/j.cub.2006.02.031.
Cooper, G.M. (2000) The Cell: A Molecular Approach. 2nd edn. Sunderland (MA): Sinauer Associates. Available at: National Center for Biotechnology Information – The Human Genome.
Kunkel, T.A. (2009). “Evolving views of DNA replication (in)fidelity.” Cold Spring Harbor Symposia on Quantitative Biology, 74, 91-101.
2. How many different ways are there to code (DNA nucleotide code) for an average human protein? In practice what are some of the reasons that all of these different codes don’t work to code for the protein of interest?
There are 61 codons that code for 20 amino acids. Most amino acids are degenerate, which means that multiple codons can code for one amino acid. The average human protein is around 400 amino acids (Milo et al., 2010). Since the average degeneracy per amino acid is approximately 3 codons, the number of possible DNA sequences is roughly 3⁴⁰⁰, which equals approximately 10¹⁹⁰ different ways to code for the same protein. some of the reasons that all of these different codes don’t work to code for the protein of interest:
- Codon usage bias (Different organisms preferentially use certain codons over others)
- mRNA secondary structure (Different codon choices create different mRNA sequences that can form stable secondary structures (hairpins, loops) that block ribosome binding or prevent translation.)
- Translation speed and protein folding (Synonymous codons translate at different speeds, and incorrect translation timing can cause the protein to misfold co-translationally, even with the correct amino acid sequence.)
References:
Milo, R., Jorgensen, P., Moran, U., Weber, G., & Springer, M. (2010). BioNumbers—the database of key numbers in molecular and cell biology. Nucleic Acids Research, 38(suppl_1), D750-D753.
Quax, T. E., Claassens, N. J., Söll, D., & van der Oost, J. (2015). Codon bias as a means to fine-tune gene expression. Molecular Cell, 59(2), 149-161.
Homework Questions from Dr. LeProust:
1. What’s the most commonly used method for oligo synthesis currently?
solid-phase phosphoramidite chemistry
2. Why is it difficult to make oligos longer than 200nt via direct synthesis?
Each time a nucleotide is added during synthesis, the coupling efficiency is about 98-99%. This means 1-2% of molecules fail to add the nucleotide at each step. As the oligo gets longer, these errors accumulate, so fewer and fewer molecules are the correct full length. By 200 nucleotides, most of the product is incomplete (truncated sequences) rather than the desired full-length oligo.
3. Why can’t you make a 2000bp gene via direct oligo synthesis?
A 2000 bp gene is too long for direct synthesis because:
- the coupling efficiency errors accumulate so much that essentially no full-length product would be made (0.99^2000 ≈ zero)
- chemical damage accumulates on the already-synthesized nucleotides during the repeated chemical cycles (especially depurination)
- the growing chain becomes physically tangled and sterically hindered on the solid support, making it harder to add new nucleotides.
References:
Guzaev, A. P. (2013). Solid‐phase supports for oligonucleotide synthesis. Current Protocols in Nucleic Acid Chemistry, 53(1), 3-1.
Kosuri, S., & Church, G. M. (2014). Large-scale de novo DNA synthesis: technologies and applications. Nature Methods, 11(5), 499-507.
Homework Question from George Church:
Choose ONE of the following three questions to answer; and please cite AI prompts or paper citations used, if any.
1. [Using Google & Prof. Church’s slide #4] What are the 10 essential amino acids in all animals and how does this affect your view of the “Lysine Contingency”?
2. [Given slides #2 & 4 (AA:NA and NA:NA codes)] What code would you suggest for AA:AA interactions?
3. [(Advanced students)] Given the one paragraph abstracts for these real 2026 grant programs sketch a response to one of them or devise one of your own:
- https://arpa-h.gov/explore-funding/programs/boss
- https://www.darpa.mil/research/programs/smart-rbc
- https://www.darpa.mil/research/programs/go
What are the 10 essential amino acids in all animals and how does this affect your view of the “Lysine Contingency”?
The 10 essential amino acids, according to Lopez et al. (2024):
- Histidine (His)
- Isoleucine (Ile)
- Leucine (Leu)
- Lysine (Lys)
- Methionine (Met)
- Phenylalanine (Phe)
- Threonine (Thr)
- Tryptophan (Trp)
- Valine (Val)
- Arginine (Arg)
Lysine, in particular, cannot be synthesised by vertebrates and must be obtained through dietary sources such as meat, dairy, eggs, and legumes (WebMD, n.d.).
The “Lysine Contingency” in Jurassic Park (both the book and the film) is presented as a genetic safety measure where the dinosaurs were engineered so they couldn’t produce the amino acid lysine, meaning they would die without supplements. But this is scientifically nonsensical, because Dinosaurs (vertibrates) already would not be capable of producing lysine on their own. Like modern animals, they would get it from their food. So there wouldn’t have been any need to genetically engineer this limitation, because it already exists in normal vertebrate biology.
References:
Jurassic Park Wiki (n.d.) Lysine contingency. Available at: https://jurassicpark.fandom.com/wiki/Lysine_contingency
Lopez, M. J., & Mohiuddin, S. S. (2024). Biochemistry, essential amino acids. In StatPearls [Internet]. StatPearls Publishing.
WebMD (n.d.) Foods High in Lysine. Available at: https://www.webmd.com/diet/foods-high-in-lysine