Week 2 HW — DNA Read, Write, and Edit

Week 2 Homework: DNA Read, Write, and Edit

In preparation for Week 2’s lecture, I reviewed the lecture slides and papers.

Professor Jacobson’s Questions

1. Polymerase Error Rate and the Human Genome

What is the error rate of polymerase?

Nature’s error-correcting polymerase has an error rate of 1 in (10^6) (1 error per 1,000,000 bases).

How does this compare to the length of the human genome?

The human genome contains approximately 3*109 base pairs (3 billion bp). If polymerase were the only fidelity mechanism, an error rate of 1:106 on a genome would result in approximately 3,000 errors every time a cell divides

How does biology deal with that discrepancy?

Biology employs Mismatch Repair systems to achieve additional error correction. Protein complex identifies and repairs mismatches that escape the polymerase’s initial proofreading, reducing the final error rate to approximately.


2. Coding Degeneracy for Human Proteins

How many different ways are there to code for an average human protein?

The average human protein is encoded by approximately 1,036 base pairs (≈345 amino acids) . Because the genetic code is redundant—with an average of roughly 3 synonymous codons per amino acid, the number of distinct DNA sequences capable of encoding the same protein sequence is approximately: 3^345

In practice, what are some reasons that all of these different codes don’t work?

While many DNA sequences encode the same amino acid sequence, they do not function equivalently in the cell due to properties of the mRNA itself

RNA Secondary Structure:
The mRNA sequence determines how the molecule folds. Certain sequences form tight secondary structures that physically block ribosome access and inhibit translation. The lecture slides showed us NUPACK simulations, showing alternative folding.

Thus, sequence matters beyond amino acid identity. The mRNA must be optimized


Dr. LeProust’s Questions

3. Current Oligo Synthesis Methods

What’s the most commonly used method for oligo synthesis currently?

The most commonly used method is solid-phase phosphoramidite chemistry. This cyclic process involves four steps repeated for each nucleotide addition:

  1. Coupling — adding the next phosphoramidite-protected nucleotide
  2. Capping — blocking unreacted 5’-OH groups to prevent truncation products
  3. Oxidation — converting the phosphite triester to a stable phosphate triester
  4. Deblocking — removing the 5’-protecting group (typically DMT) to prepare for the next cycle

4. Length Limitations in Direct Synthesis

Why is it difficult to make oligos longer than 200nt via direct synthesis?

The difficulty arises from the accumulation of truncation products and synthesis errors over many cycles . Each coupling step has a typical efficiency of 98-99.5%. For a 200-nucleotide oligo: 16%

As length increases, the proportion of full-length product drops exponentially, while incomplete sequences (n-1, n-2, etc.) accumulate. The slides illustrated this with chromatograms showing significant impurity peaks for 500-nucleotide synthesis attempts .

Additionally, depurination and other side reactions become more probable with longer sequences, further reducing fidelity.


5. Gene Assembly vs. Direct Synthesis

Why can’t you make a 2000bp gene via direct oligo synthesis?

Direct phosphoramidite synthesis is ~200-500 nucleotides due to error accumulation and yield collapse . To construct a 2000bp gene, we ca, use gene assembly methods:

  1. Synthesize many shorter oligonucleotides (typically 40-200mers)
  2. Design overlapping sequences between adjacent oligos
  3. Assemble via PCR (polymerase chain reaction) or enzymatic assembly (e.g., Gibson assembly, Golden Gate)

This modular approach allows the cumulative error rate to be managed across multiple synthesis reactions, with final assembly and error correction performed enzymatically .


George Church’s Question (Option 1)

6. Essential Amino Acids and the “Lysine Contingency”

What are the 10 essential amino acids in all animals?

The 10 amino acids considered essential for animals are :

  1. Arginine (R)
  2. Histidine (H)
  3. Isoleucine (I)
  4. Leucine (L)
  5. Lysine (K)
  6. Methionine (M)
  7. Phenylalanine (F)
  8. Threonine (T)
  9. Tryptophan (W)
  10. Valine (V)

How does this affect your view of the “Lysine Contingency”?

The “Lysine Contingency” from Jurassic Park proposed engineering dinosaurs unable to produce lysine, ensuring they would die without dietary supplements provided by the park. Understanding that lysine is an essential amino acid reveals this contingency as scientifically flawed for two reasons:

By definition, an “essential” amino acid is one that vertebrates already cannot synthesize. Animals naturally lack the lysine biosynthesis pathway (found in bacteria, plants, and fungi). Therefore:

  • The dinosaurs would already be natural lysine auxotrophs without any genetic engineering
  • “Removing” the ability to produce lysine is impossible—the pathway was never present in vertebrate genomes
  • The genetic modification described in the film would be unnecessary or meaningless
Note
If receptors exceed this length, I'll need to use gene assembly approaches rather than ordering full-length oligos—affecting both cost and timeline for Aim 1.