Week 2 HW: DNA Read Write and Edit

Part 1: Benchling & In-silico Gel Art

  • Make a free account at benchling.com
  • Import the Lambda DNA.
  • Simulate Restriction Enzyme Digestion with the following Enzymes: EcoRI, HindIII, BamHI, KpnI, EcoRV, SacI, SalI cover image cover image
  • Create a pattern/image in the style of Paul Vanouse’s Latent Figure Protocol artworks.
image image

Image: modified in photoshop

image image

Image: modified in photoshop

image image

Image: modified in photoshop

ComponentM1234567891011
Water19 µL14 µL13 µL13 µL13 µL13 µL14 µL13 µL13 µL13 µL13 µL14 µL
CutSmart Buffer-2 µL2 µL2 µL2 µL2 µL2 µL2 µL2 µL2 µL2 µL2 µL
λ DNA-3 µL3 µL3 µL3 µL3 µL3 µL3 µL3 µL3 µL3 µL3 µL
Enzyme(s)1 µL Ladder1 µL SacI1 µL EcoRI1 µL SacI1 µL KpnI, 1 µL SalI1 µL KpnI, 1 µL BamHI1 µL HindIII1 µL KpnI, 1 µL BamHI1 µL KpnI, 1 µL SalI1 µL SacI, 1 µL SalI1 µL EcoRI, 1 µL SalI1 µL SacI

Part 3: DNA Design Challenge

3.1. Choose a protein that you find interesting. Which protein have you chosen and why? Using one of the tools described in recitation (NCBI, UniProt, google), obtain the protein sequence for the protein you chose.

[Example from our group homework, you may notice the particular format — The example below came from UniProt]

>sp|P03609|LYS_BPMS2 Lysis protein OS=Escherichia phage MS2 OX=12022 PE=2 SV=1 METRFPQQSQQTPASTNRRRPFKHEDYPCRRQQRSSTLYVLIFLAIFLSKFTNQLLLSLL EAVIRTVTTLQQLLT

Proteins research:

  • Aquaporins - facilitates transport of water, glycerol, and other small solutes across biological membranes
  • Hydrophobins - modulates hydrophobicity to aid aerial growth and environmental adaptation.
  • GTPase - molecular switches

The protein I chose is VMH3-1, a Class I hydrophobin present in Pleurotus ostreatus strain PC15. This forms a hydrophobic coating on the cell walls allowing mycelium to grow in damp environments without sinking into them for example: soil

[tr|Q8WZI4|Q8WZI4_PLEOS Hydrophobin OS=Pleurotus ostreatus OX=5322 GN=vmh3-1 PE=3 SV=1 MFFQTTIVAALASLAVATPLALRTDSRCNTESVKCCNKSEDAETFKKSASAALIPIKIGD ITGKVYSECSPIVGLIGGSSCSAQTVCCDNAKFNGLVNIGCTPINVAL]

https://www.uniprot.org/uniprotkb/Q8WZI4/entry

3.2. Reverse Translate: Protein (amino acid) sequence to DNA (nucleotide) sequence.

The Central Dogma discussed in class and recitation describes the process in which DNA sequence becomes transcribed and translated into protein. The Central Dogma gives us the framework to work backwards from a given protein sequence and infer the DNA sequence that the protein is derived from. Using one of the tools discussed in class, NCBI or online tools (google “reverse translation tools”), determine the nucleotide sequence that corresponds to the protein sequence you chose above.

[Example: Get to the original sequence of phage MS2 L-protein from its genome phage MS2 genome - Nucleotide - NCBI]

Lysis protein DNA sequence atggaaacccgattccctcagcaatcgcagcaaactccggcatctactaatagacgccggccattcaaacatgaggattacccatgtcgaagacaacaaagaagttcaactctttatgtattgatcttcctcgcgatctttctctcgaaatttaccaatcaattgcttctgtcgctactggaagcggtgatccgcacagtgacgactttacagcaattgcttacttaa

VMH3-1 protein Nucleotide sequence [>AJ420971.1 Pleurotus ostreatus vmh3-1 gene for hydrophobin 3 (allele 1), exons 1-3 ATGTTCTTCCAAACTACCATCGTCGCCGCCCTCGCTTCCCTTGCGGTCGCCACTCCTCTCGCACTTCGCA CTGACAGTCGCTGCAACACCGAGTCCGTGAAGTGCTGCAACAAGTCTGAGGATGCAGAGACCTTCAAGAA GAGCGCGTCGGCCGCCCTCATCCCGATTAAGATCGGTGATATTACCGGCAAGGTGTACTCGGAGTGTTCT CCCATTGTCGGCCTCATTGGCGGGTCTAGCTGGTACGTGTCTTTGTGCGTCTCTGATGTCAAGTCTGTTC TGACTCTTTTTTCAGCTCCGCGCAAACCGTTTGCTGCGATAACGCTAAATTCAGTAAGCAATCATTCTTG GGCCTCTTCATTGACTTTCGGCGGAGAATTTGGTACTAATTCTTCCGCATGTTAGATGGTCTCGTCAACA TTGGATGCACGCCCATCAACGTTGCCTTGTAA]

https://www.ncbi.nlm.nih.gov/nuccore/AJ420971.1?report=fasta

3.3. Codon optimization

Once a nucleotide sequence of your protein is determined, you need to codon optimize your sequence. You may, once again, utilize google for a “codon optimization tool”. In your own words, describe why you need to optimize codon usage. Which organism have you chosen to optimize the codon sequence for and why?

[Example from Codon Optimization Tool | Twist Bioscience while avoiding Type IIs enzyme recognition sites BsaI, BsmBI, and BbsI]

Lysis protein DNA sequence with Codon-Optimization ATGGAAACCCGCTTTCCGCAGCAGAGCCAGCAGACCCCGGCGAGCACCAACCGCCGCCGCCCGTTCAAACATGAAGATTATCCGTGCCGTCGTCAGCAGCGCAGCAGCACCCTGTATGTGCTGATTTTTCTGGCGATTTTTCTGAGCAAATTCACCAACCAGCTGCTGCTGAGCCTGCTGGAAGCGGTGATTCGCACAGTGACGACCCTGCAGCAGCTGCTGACCTAA

image image

https://eu.idtdna.com/CodonOpt

I chose Escherichia coli or E.coli for codon optimisation sequence due to:

  • Ease of replication: E. coli doubles every 20-30 min (fungi = days)
  • High production: 100-500 mg/L vs 1-10 mg/L native PC15
  • Proven system: Industry standard for recombinant hydrophobins

Optimising codon usage is important because not all codons are used equally in every organism. I learnt that even though multiple codons can code for the same amino acid, some are preferred over others depending on the species. (While I was using IDT Codon optimisation tool, I checked for different organims) Using codons that match the host organism’s preferred usage can improve translation efficiency and protein expression (Mäkelä et al., 2020).

If codons are rare in the host, the ribosome may stall during translation due to limited availability of the corresponding tRNAs, leading to slower protein synthesis, misfolded proteins, or low yield. Optimizing codons ensures smoother, faster translation and can enhance protein stability, functional folding, and overall biotechnological productivity (Mäkelä et al., 2020).

One detail I learnt (probably insignificant for science specialists): I initially tried optimising from the DNA sequence directly, but it did not give accurate results because the sequence could not be properly sorted into triplets. I attempted decoding it using perplexity, but that also failed to produce an answer. Through the chats, I learnt that the DNA sequence could be written differntly obstained from translation of the amino acid (AA) sequence. I recalled how Prof. George mentioned in class that different amino acids have different codons. I then tried optimising the sequence starting from the AA sequence from UniProt and was able to successfully generate an optimised codon sequence for E. coli.

Reference: Mäkelä, M. et al. (2020) ‘Codon optimization with deep learning to enhance protein expression’, Scientific Reports, 10, p. 16968. Available at: https://www.nature.com/articles/s41598-020-74091-z

3.4. You have a sequence! Now what?

What technologies could be used to produce this protein from your DNA? Describe in your words the DNA sequence can be transcribed and translated into your protein. You may describe either cell-dependent or cell-free methods, or both.

DNA READ

(i) What DNA would you want to sequence (e.g., read) and why? This could be DNA related to human health (e.g. genes related to disease research), environmental monitoring (e.g., sewage waste water, biodiversity analysis), and beyond (e.g. DNA data storage, biobank).

(ii) In lecture, a variety of sequencing technologies were mentioned. What technology or technologies would you use to perform sequencing on your DNA and why? Also answer the following questions:

1. Is your method first-, second- or third-generation or other? How so?

2. What is your input? How do you prepare your input (e.g. fragmentation, adapter ligation, PCR)? List the essential steps.

3. What are the essential steps of your chosen sequencing technology, how does it decode the bases of your DNA sample (base calling)?

4. What is the output of your chosen sequencing technology?

DNA WRITE

(i) What DNA would you want to synthesize (e.g., write) and why?

These could be individual genes, clusters of genes or genetic circuits, whole genomes, and beyond. As described in class thus far, applications could range from therapeutics and drug discovery (e.g., mRNA vaccines and therapies) to novel biomaterials (e.g. structural proteins), to sensors (e.g., genetic circuits for sensing and responding to inflammation, environmental stimuli, etc.), to art (DNA origamis). If possible, include the specific genetic sequence(s) of what you would like to synthesize!

(ii) What technology or technologies would you use to perform this DNA synthesis and why?

Also answer the following questions:

1. What are the essential steps of your chosen sequencing methods?

2. What are the limitations of your sequencing method (if any) in terms of speed, accuracy, scalability?

DNA EDIT

(i) What DNA would you want to edit and why?

In class, George shared a variety of ways to edit the genes and genomes of humans and other organisms. Such DNA editing technologies have profound implications for human health, development, and even human longevity and human augmentation. DNA editing is also already commonly leveraged for flora and fauna, for example in nature conservation efforts, (animal/plant restoration, de-extinction), or in agriculture (e.g. plant breeding, nitrogen fixation). What kinds of edits might you want to make to DNA (e.g., human genomes and beyond) and why?

(ii) What technology or technologies would you use to perform these DNA edits and why?

Also answer the following questions:

1. How does your technology of choice edit DNA? What are the essential steps?

2. What preparation do you need to do (e.g. design steps) and what is the input (e.g. DNA template, enzymes, plasmids, primers, guides, cells) for the editing?

3. What are the limitations of your editing methods (if any) in terms of efficiency or precision?