DNA Structure and Replication - Complete Interactive Lesson
Part 1: DNA Structure
DNA Structure — The Molecular Basis of Heredity
Part 1 of 7
Understanding DNA replication requires a thorough understanding of DNA structure. The double helix model, proposed by Watson and Crick in 1953 (based on X-ray crystallography data from Rosalind Franklin and Maurice Wilkins), remains one of the most important discoveries in biology.
Nucleotide Structure
DNA is a polymer of nucleotides. Each nucleotide has three components:
- Deoxyribose sugar (5-carbon sugar lacking an -OH group at the 2' carbon)
- Phosphate group (attached to the 5' carbon of the sugar)
- Nitrogenous base (attached to the 1' carbon of the sugar)
The four DNA bases:
| Base | Type | Rings | Pairs with |
|---|---|---|---|
| Adenine (A) | Purine | 2 rings | Thymine (T) — 2 hydrogen bonds |
| Guanine (G) | Purine | 2 rings | Cytosine (C) — 3 hydrogen bonds |
| Thymine (T) | Pyrimidine | 1 ring | Adenine (A) — 2 hydrogen bonds |
| Cytosine (C) | Pyrimidine | 1 ring | Guanine (G) — 3 hydrogen bonds |
Chargaff's Rules: In any DNA molecule, the amount of A equals the amount of T, and the amount of G equals the amount of C. This 1:1 ratio (A=T, G=C) was critical evidence for the base-pairing model.
Key structural features of the double helix:
- Two antiparallel strands (one runs 5' → 3', the other 3' → 5')
- Sugar-phosphate backbones on the outside; bases face inward
- Bases held together by hydrogen bonds (A-T: 2 H-bonds; G-C: 3 H-bonds)
- The helix has a major groove and a minor groove (important for protein-DNA interactions)
- One complete turn = 10 base pairs = 3.4 nm
- Diameter = 2 nm
Checkpoint — DNA Structure
DNA Packaging in Eukaryotes
Human cells contain about 6.4 billion base pairs of DNA (~2 meters per cell if stretched out). This must be packaged into a nucleus only ~6 m in diameter — a compaction ratio of ~10,000:1.
Packaging hierarchy:
- Nucleosome — 147 bp of DNA wraps ~1.65 times around a histone octamer (2 copies each of H2A, H2B, H3, H4); the fundamental unit of chromatin (11 nm "beads on a string")
- Linker histone H1 binds between nucleosomes, helping them pack into a 30 nm fiber
- Looped domains — 30 nm fiber forms loops of 30,000-100,000 bp attached to a protein scaffold
- Heterochromatin — maximally condensed, transcriptionally inactive
- Metaphase chromosome — highest compaction (~10,000-fold), visible under a light microscope
Epigenetics and Histones: Chemical modifications to histones (acetylation, methylation, phosphorylation) alter chromatin structure and gene accessibility. Histone acetylation (by HATs) generally OPENS chromatin (euchromatin, active). Histone deacetylation (by HDACs) CLOSES it (heterochromatin, silenced).
Key Terms — DNA Structure
Exit Ticket — DNA Structure
Part 2: Semiconservative Replication
Semiconservative Replication
Part 2 of 7
Three models for DNA replication were proposed:
- Conservative — the original double helix remains intact; a completely new copy is made
- Semiconservative — each new double helix consists of one original (parental) strand and one new strand
- Dispersive — parental and new DNA are interspersed throughout both strands
The Meselson-Stahl experiment (1958) elegantly determined which model is correct.
The Meselson-Stahl Experiment
Design:
- E. coli were grown for many generations in medium containing heavy nitrogen (N) — all DNA became uniformly "heavy"
- Cells were transferred to medium containing light nitrogen (N)
- DNA was extracted after each generation and centrifuged in a CsCl density gradient
Results:
| Generation | DNA bands observed | Interpretation |
|---|---|---|
| 0 (all N) | One heavy band | All DNA is heavy |
| 1 (one round of replication in N) | One intermediate band | Each DNA molecule has one heavy and one light strand |
| 2 | Half intermediate, half light | Half the molecules retain a heavy strand; half are entirely light |
| 3 | 1/4 intermediate, 3/4 light | Pattern continues predictably |
Conclusion: DNA replication is semiconservative — each daughter molecule contains one parental strand and one newly synthesized strand.
Why this rules out other models:
- Conservative would show heavy + light bands at generation 1 (not intermediate)
- Dispersive would show all intermediate bands that become progressively lighter (never a pure light band)
- Only semiconservative predicts one intermediate band at generation 1, then both intermediate and light at generation 2
Checkpoint — Semiconservative Replication
Origins of Replication
DNA replication begins at specific sequences called origins of replication:
Prokaryotes:
- Single circular chromosome with one origin (oriC in E. coli)
- Replication proceeds bidirectionally from the origin, creating two replication forks that move in opposite directions and meet on the opposite side of the chromosome
- Total replication time: ~40 minutes
Eukaryotes:
- Multiple linear chromosomes with many origins (30,000-50,000 in human cells)
- Multiple origins allow the much larger genome to be replicated within the S phase time window
- Each origin fires once per S phase (controlled by the licensing system — pre-replication complexes mark origins for use)
- Adjacent origins define a replicon — the segment of DNA replicated from one origin
- Replication forks from adjacent origins meet and the replicons fuse
Why multiple origins? Human DNA polymerase moves at ~50 nucleotides/second. With 6.4 billion bp and bidirectional replication from one origin, it would take ~2 years. Multiple origins reduce this to ~8 hours.
Key Terms — Replication Model
Exit Ticket
Part 3: Enzymes of Replication
The Replication Machinery
Part 3 of 7
DNA replication requires a team of enzymes and proteins working together at the replication fork — the Y-shaped region where the parental double helix is unwound and new strands are synthesized.
Key Enzymes and Proteins
| Enzyme/Protein | Function |
|---|---|
| Helicase | Unwinds the double helix by breaking hydrogen bonds between bases; moves along the DNA using ATP hydrolysis |
| Single-strand binding proteins (SSB) | Bind to exposed single-stranded DNA to prevent reannealing and protect from nuclease degradation |
| Topoisomerase (Gyrase in prokaryotes) | Relieves torsional strain (supercoiling) ahead of the replication fork by cutting, swiveling, and re-joining the DNA backbone |
| Primase | RNA polymerase that synthesizes a short RNA primer (5-10 nucleotides) complementary to the template; provides the free 3'-OH needed by DNA polymerase |
| DNA Polymerase III (prokaryotes) | The main replicative polymerase; adds nucleotides to the 3' end of the primer/growing strand in the 5' → 3' direction |
| DNA Polymerase I (prokaryotes) | Removes RNA primers (5' → 3' exonuclease activity) and replaces them with DNA |
| DNA Ligase | Seals the nick (phosphodiester bond) between adjacent Okazaki fragments after primer removal |
| Sliding clamp (PCNA in eukaryotes) | Ring-shaped protein that encircles DNA and tethers DNA polymerase to the template, increasing processivity |
| Clamp loader | Uses ATP to load the sliding clamp onto DNA at primer-template junctions |
Why primers? DNA polymerase can only ADD nucleotides to an existing 3'-OH group. It cannot start a new strand from scratch. Primase (an RNA polymerase) CAN start de novo, providing the initial 3'-OH.
Checkpoint — Replication Enzymes
Leading and Lagging Strands
Because DNA polymerase can only synthesize in the 5' → 3' direction, and the two parental strands are antiparallel, the two daughter strands are synthesized differently:
Leading strand:
- Template runs 3' → 5' (so new strand grows 5' → 3' toward the fork)
- Synthesized continuously with a single primer
- DNA polymerase moves in the same direction as the fork
Lagging strand:
- Template runs 5' → 3' (so new strand must grow 5' → 3' AWAY from the fork)
- Synthesized discontinuously as a series of short fragments called Okazaki fragments
- ~1000-2000 nt in prokaryotes; ~100-200 nt in eukaryotes
- Each Okazaki fragment requires its own RNA primer
- After synthesis, DNA Polymerase I removes the RNA primers and fills in with DNA
- DNA Ligase seals the remaining nicks, joining the Okazaki fragments into a continuous strand
The lagging strand is the "complex" strand — it requires repeated cycles of priming, extension, primer removal, gap filling, and ligation. This makes lagging strand synthesis more error-prone and slower than leading strand synthesis.
Checkpoint — Leading and Lagging
Key Terms — Replication Machinery
Exit Ticket
Part 4: Leading vs Lagging Strand
Proofreading and DNA Repair
Part 4 of 7
DNA replication must be extraordinarily accurate — the error rate is approximately 1 mistake per 10 base pairs in E. coli. This remarkable fidelity is achieved through three layers of error correction.
Three Layers of Replication Fidelity
Layer 1: Base selection by DNA polymerase (~10 accuracy)
- DNA polymerase has a tight active site that favors correct Watson-Crick base pairs (A-T, G-C)
- Incorrect bases fit poorly and are rejected before incorporation
- Error rate: ~1 in 100,000
Layer 2: Proofreading (3' → 5' exonuclease activity, ~10 improvement)
- DNA polymerase has a built-in editor: if a wrong nucleotide is incorporated, the polymerase detects the mismatch (distortion in the helix)
- The polymerase reverses direction and removes the incorrect nucleotide using 3' → 5' exonuclease activity
- A correct nucleotide is then inserted
- Combined error rate: ~1 in 10
Layer 3: Mismatch repair (MMR, ~10 improvement)
- After replication, mismatch repair proteins scan the newly synthesized DNA
- They detect and correct remaining mismatches
- The key challenge: distinguishing which strand has the error (old vs. new strand)
- In E. coli: the parental strand is methylated (GATC sites); the new strand is not yet methylated, so repair enzymes know to fix the new strand
- In eukaryotes: the new strand is identified by the presence of nicks (gaps not yet sealed)
- Combined final error rate: ~1 in 10
Checkpoint — Proofreading and Repair
DNA Damage and Additional Repair Mechanisms
Beyond replication errors, DNA is constantly damaged by environmental and metabolic factors:
Types of DNA damage:
- Deamination: Spontaneous loss of an amino group from cytosine, converting it to uracil (if not repaired, G-C becomes A-T after replication)
- Depurination: Loss of a purine base (A or G) from the sugar-phosphate backbone (~5000 per cell per day)
- Thymine dimers: UV light causes adjacent thymines to form covalent bonds (pyrimidine dimers), distorting the helix
- Oxidative damage: Reactive oxygen species (ROS) modify bases (e.g., 8-oxoguanine mispairs with adenine)
- Alkylation: Chemical agents add alkyl groups to bases
Repair pathways:
| Pathway | Damage type | Mechanism |
|---|---|---|
| Base excision repair (BER) | Modified/damaged single bases | Glycosylase removes damaged base → AP endonuclease cuts backbone → polymerase fills gap → ligase seals |
| Nucleotide excision repair (NER) | Bulky lesions (thymine dimers, crosslinks) | Endonucleases cut on both sides of damage → ~12 nt oligonucleotide removed → polymerase fills → ligase seals |
| Homologous recombination (HR) | Double-strand breaks (DSBs) | Uses sister chromatid as template for accurate repair — only in S/G phase |
| Non-homologous end joining (NHEJ) | Double-strand breaks | Directly ligates broken ends — faster but error-prone (may lose bases) |
Xeroderma pigmentosum (XP): A genetic disease caused by mutations in NER genes. Patients cannot repair UV-induced thymine dimers and are extremely sensitive to sunlight, with high rates of skin cancer.
Checkpoint — DNA Repair
Key Terms — DNA Repair
Exit Ticket
Part 5: Proofreading & Repair
Telomeres and the End Replication Problem
Part 5 of 7
Linear chromosomes in eukaryotes face a unique challenge: the end replication problem. This problem does not exist in prokaryotes because their chromosomes are circular.
The End Replication Problem
The problem:
- On the lagging strand, an RNA primer must initiate each Okazaki fragment
- At the very end of the chromosome (3' end of the template), the last RNA primer is synthesized, and DNA polymerase extends from it
- When this primer is removed, there is a short gap at the 5' end of the new strand that CANNOT be filled — there is no upstream 3'-OH for DNA polymerase to extend from
- Result: the daughter strand is slightly shorter than the parent
Consequence: With each round of replication, chromosomes get shorter at both ends. After many divisions, essential genes near the ends would be lost.
Telomeres — The Protective Solution
Telomeres are repetitive, non-coding DNA sequences at the ends of linear chromosomes:
- Human telomere repeat: TTAGGG (repeated 1000-2000 times, totaling 5-15 kb)
- Telomeres provide a "buffer zone" of expendable sequence — shortening removes repeats, not genes
- Telomeres also form a protective structure called a T-loop (the 3' overhang folds back and invades the double-stranded region) with a protein complex called shelterin that prevents the cell from recognizing chromosome ends as DNA breaks
Hayflick Limit: Normal somatic cells can divide approximately 50-70 times before telomeres become critically short. At this point, cells enter replicative senescence (a permanent G state) or undergo apoptosis. This is a tumor-suppression mechanism.
Checkpoint
Telomerase — Extending the Ends
Telomerase is a specialized enzyme that extends telomeres, counteracting the end replication problem:
Structure:
- Telomerase is a ribonucleoprotein (protein + RNA)
- Contains TERT (telomerase reverse transcriptase) — the catalytic protein subunit
- Contains TERC (telomerase RNA component) — includes a template sequence complementary to the telomeric repeat
Mechanism:
- The TERC template (3'-AAUCCC-5') base-pairs with the 3' overhang of the telomere
- TERT extends the 3' end using the RNA template (reverse transcription — RNA → DNA)
- Telomerase translocates and repeats, adding multiple TTAGGG repeats
- Primase then synthesizes a primer on the extended 3' overhang
- DNA polymerase fills in the complementary strand
- The primer is removed, leaving a slightly extended chromosome
Telomerase expression:
- Active in: germ cells, stem cells, early embryonic cells — these cells must divide indefinitely
- Inactive in: most somatic cells — contributes to the Hayflick limit and aging
- Reactivated in: ~85-90% of cancers — telomerase reactivation grants immortality
Cancer Connection: Telomerase reactivation is one of the hallmarks of cancer. Drugs targeting telomerase (e.g., imetelstat) are being developed as potential cancer therapies. However, targeting telomerase could also affect stem cells, posing a therapeutic challenge.
Checkpoint — Telomerase
Key Terms — Telomeres
Exit Ticket
Part 6: Problem-Solving Workshop
Problem-Solving Workshop — DNA Replication
Part 6 of 7
This workshop applies DNA replication concepts to experimental scenarios and quantitative problems.
Scenario 1: Replication Fork Analysis
A researcher treats E. coli with radioactive thymidine (H-thymidine) for a brief pulse, then chases with unlabeled thymidine. After autoradiography of the replicating DNA:
- The label appears as a band along the newly synthesized DNA
- The leading strand shows a continuous band of label
- The lagging strand shows a series of short labeled segments (Okazaki fragments) with gaps where primers were located
If the pulse is very short (seconds), only the most recently synthesized DNA is labeled. The leading strand shows label near the fork, while the lagging strand shows label in the most recently completed Okazaki fragment.
If the chase is long enough, DNA Pol I replaces primers with DNA and ligase joins fragments, so the lagging strand eventually looks continuous.
Scenario 1 Questions
Scenario 2: Density Gradient Predictions
Starting with one double-stranded DNA molecule where BOTH strands are labeled with N (heavy):
After 1 generation in N:
- 2 molecules, each with one N strand + one N strand = 2 intermediate
After 2 generations in N:
- 4 molecules total
- 2 have one N + one N = 2 intermediate
- 2 have both N = 2 light
After n generations:
- Total molecules =
- Intermediate molecules = always 2 (the two original parental strands + a new partner)
- Light molecules =
- Heavy molecules = 0 (after generation 1)
Quantitative AP Tip: The number of intermediate-density molecules never changes (always 2) because the two original heavy strands are conserved indefinitely, each paired with a new light strand.
Scenario 2 Questions
Apply Your Knowledge
Exit Ticket — Workshop
Part 7: AP Review
AP Review — DNA Replication
Part 7 of 7
Comprehensive AP-exam-style questions integrating all DNA replication concepts.
Key Principles Summary
- DNA replication is semiconservative — each daughter molecule has one parental and one new strand (Meselson-Stahl)
- Replication is bidirectional from origins and proceeds in the 5' → 3' direction only
- The leading strand is continuous; the lagging strand is discontinuous (Okazaki fragments)
- Primase provides RNA primers; DNA Polymerase III extends; Pol I removes primers; Ligase seals nicks
- Three layers of fidelity: base selection, proofreading (3'→5' exonuclease), and mismatch repair
- Multiple DNA repair pathways (BER, NER, HR, NHEJ) protect against different types of damage
- Telomeres protect chromosome ends; telomerase counteracts the end replication problem