🎯 Key Points
- DNA: antiparallel double helix, A-T (2 H-bonds), G-C (3 H-bonds); one turn = 3.4nm = 10 bp; replication is semi-conservative (Meselson-Stahl)
- Eukaryotic DNA packaging: DNA + histone octamer (2 each of H2A,H2B,H3,H4) = nucleosome (~200bp); euchromatin = loose/active, heterochromatin = dense/inactive
- Transcription unit has a promoter, structural gene, and terminator; only ONE strand (template/antisense) is read 3'→5', mRNA built 5'→3'
- Genetic code: triplet, degenerate (multiple codons per amino acid), nearly universal, AUG = start (also codes Met), UAA/UAG/UGA = stop
- Lac operon: structural genes OFF by default (repressor bound); lactose (inducer) binds repressor, releasing it, switching genes ON — classic negative regulation example
- Human Genome Project: ~3 billion bp, only ~20,000-25,000 genes, >90% non-coding DNA
Helicase unwinds the double helix at the replication fork; one new strand (leading) is synthesised continuously toward the fork, while the other (lagging) must be made in short, separate Okazaki fragments because DNA polymerase only adds nucleotides 5'→3'.
DNA Structure: The Double Helix
- Watson and Crick (1953) proposed the double helix model of DNA, based on X-ray diffraction data generated by Rosalind Franklin and Maurice Wilkins.
- DNA is made of two polynucleotide chains coiled in a right-handed double helix; the two strands are antiparallel (one runs 5' to 3', the other 3' to 5').
- Each nucleotide has three parts: a nitrogenous base, a deoxyribose sugar, and a phosphate group. Sugar and phosphate form the backbone; bases project inward.
- Base pairing rules (complementarity): Adenine (A) pairs with Thymine (T) via 2 hydrogen bonds; Guanine (G) pairs with Cytosine (C) via 3 hydrogen bonds. This is why G-C rich DNA is more thermally stable.
- Chargaff's rule: in any double-stranded DNA, the amount of A equals T, and the amount of G equals C, so purines (A+G) equal pyrimidines (T+C).
- The two strands, if separated, would each act as a template for synthesis of a complementary strand; this property is the structural basis of replication.
- One full turn of the helix spans about 3.4 nm and contains 10 base pairs; consecutive bases are 0.34 nm apart; the diameter of the helix is about 2 nm.
Packaging of DNA: Nucleosomes
- The length of DNA in a human cell is about 2.2 metres but it is packaged inside a nucleus only a few micrometres in diameter, achieved through extensive coiling and protein packaging.
- In prokaryotes (e.g., E. coli) negatively charged DNA is held in "looped domains" with positively charged proteins in a region called the nucleoid.
- In eukaryotes, DNA wraps around a core of eight histone proteins (two molecules each of H2A, H2B, H3, and H4) to form a nucleosome. Histones are rich in basic amino acids (lysine and arginine), giving them a net positive charge that binds the negatively charged DNA.
- A typical nucleosome contains about 200 bp of DNA helix; nucleosomes are repeating units of chromatin, seen as "beads on a string" under the electron microscope.
- H1 histone is found at the linker DNA between nucleosomes and helps in further packaging into higher-order chromatin fibres.
- Chromatin that is loosely packed and stains light is called euchromatin (transcriptionally active); densely packed, dark-staining chromatin is heterochromatin (transcriptionally inactive).
DNA Replication
- DNA replicates semiconservatively: each daughter DNA molecule has one parental (old) strand and one newly synthesised strand.
- Meselson and Stahl's experiment (1958) confirmed this using E. coli grown in a medium with heavy nitrogen (15N), then transferred to normal 14N medium. Density-gradient (CsCl) centrifugation of DNA at successive generations showed hybrid (15N-14N) and light (14N-14N) DNA in the ratios predicted only by the semiconservative model.
- Replication begins at a specific sequence called origin of replication; in eukaryotic chromosomes there are multiple origins.
- Helicase unwinds the parental double helix at the replication fork, creating two template strands.
- Since DNA polymerase can only add nucleotides to an existing 3'-OH end, a short RNA stretch called a primer, synthesised by primase, is required to initiate synthesis.
- DNA polymerase catalyses polymerisation only in the 5' to 3' direction. On the template strand that runs 3' to 5' (relative to the new strand direction), synthesis is continuous, forming the leading strand. On the other template, synthesis is discontinuous, producing short fragments called Okazaki fragments, which together form the lagging strand.
- DNA ligase joins the Okazaki fragments by sealing the nicks (forming phosphodiester bonds) into a continuous lagging strand; RNA primers are removed and replaced with DNA before ligation.
- Replication occurs during the S-phase of the cell cycle, in association with the duplication of centriole and other organelles.
- The overall reaction is energetically favourable because the deoxynucleoside triphosphates (dNTPs) used as substrates are high-energy compounds that provide energy for both bond formation and polymerisation.

The replication fork: continuous synthesis on the leading strand, Okazaki fragments on the lagging strand. Image: LadyofHats (Mariana Ruiz), Public Domain, via Wikimedia Commons.
The Central Dogma and Transcription
- The central dogma (Francis Crick) states that genetic information flows from DNA to RNA to protein: DNA → RNA → Protein.
- Transcription is the synthesis of an RNA copy from a DNA template, catalysed by RNA polymerase. Only a segment of DNA and only one of its two strands (the template strand) is transcribed.
- The strand that is used as a template is called the template strand; the other strand, which has the same sequence as the RNA transcript (except T replaced by U), is called the coding strand.
- Transcription requires a promoter sequence (where RNA polymerase binds and initiation begins) and a terminator sequence (defines the end of transcription); both flank the structural gene as recognition sequences for RNA polymerase.
- In bacteria, a single RNA polymerase catalyses initiation, elongation, and termination; a transcription factor, sigma, helps initiate, and another, rho, is involved in termination.
- In eukaryotes there are three RNA polymerases: RNA polymerase I transcribes rRNA, RNA polymerase III transcribes tRNA and small RNAs, and RNA polymerase II transcribes the precursor of mRNA (hnRNA).
- Post-transcriptional processing (in eukaryotes): the primary transcript (hnRNA) undergoes splicing (removal of non-coding introns and joining of coding exons), capping (addition of methyl guanosine triphosphate cap at the 5' end), and tailing (addition of poly-A tail, about 200-300 adenylate residues, at the 3' end) to form mature mRNA.

Transcription: RNA polymerase reads the template strand 3→5 and builds RNA 5→3. Public Domain, via Wikimedia Commons.
The Genetic Code
- The genetic code is the relationship between the sequence of nucleotides in mRNA (codons) and the sequence of amino acids in the protein. It was largely deciphered by Har Gobind Khorana and confirmed biochemically by Marshall Nirenberg; Severo Ochoa helped synthesise an enzyme (polynucleotide phosphorylase) used to polymerise RNA with defined sequences.
- A codon is a triplet of nucleotides; there are 4^3 = 64 codons in total.
- AUG codes for methionine and also functions as the universal start codon.
- Three codons, UAA, UAG, and UGA, do not code for any amino acid and act as stop (termination) codons.
- Degeneracy: the code is degenerate, meaning most amino acids are coded by more than one codon (e.g., leucine and serine have 6 codons each); only methionine and tryptophan have a single codon each.
- The code is nearly universal: from bacteria to humans, UUU codes for phenylalanine in almost all organisms (with minor exceptions in some organelles and protozoans).
- The code is unambiguous: a given codon always codes for the same amino acid; it is also non-overlapping and comma-less, read continuously in triplets without gaps.
- Wobble hypothesis (Francis Crick): the third (3') base of the codon often pairs less strictly with the corresponding first base of the tRNA anticodon, allowing a single tRNA to recognise more than one codon for the same amino acid, explaining why far fewer than 61 tRNAs are needed to read all sense codons.
tRNA: The Adaptor Molecule
- Francis Crick postulated the existence of an adaptor molecule that would read the genetic code and bind specific amino acids; this is the tRNA (transfer RNA).
- tRNA has a characteristic cloverleaf secondary structure, which folds further into an inverted L-shaped (or "L"-folded) tertiary structure.
- Each tRNA has an anticodon loop with bases complementary to a specific mRNA codon, and an acceptor end (3' end) to which a specific amino acid is attached (charged by aminoacyl-tRNA synthetase, forming aminoacyl-tRNA).
- A separate initiator tRNA recognises the start codon AUG to begin translation.
Translation
- Translation is the polymerisation of amino acids to form a polypeptide chain, with the sequence and order of amino acids defined by the sequence of codons in the mRNA, occurring at the ribosome.
- Ribosomes are the actual sites of protein synthesis; they consist of two subunits that remain dissociated when not translating, and associate around an mRNA when active. The ribosome also acts as a catalyst (23S rRNA in bacteria acts as a ribozyme) for the formation of peptide bonds.
- Stages: initiation (small subunit binds mRNA at start codon, initiator tRNA brings methionine), elongation (ribosome facilitates entry of aminoacyl-tRNAs, forms peptide bonds, and translocates along mRNA in the 5' to 3' direction), and termination (a release factor binds at the stop codon, terminating translation and releasing the polypeptide).

Translation: tRNA delivers amino acids to the A site; the polypeptide grows from the P site. Image: LadyofHats, Public Domain, via Wikimedia Commons.
Regulation of Gene Expression: The Lac Operon
- Gene expression in prokaryotes is commonly regulated at the level of transcription initiation; an operon is a cluster of structural genes under the control of a single promoter, transcribed as one polycistronic mRNA.
- The lac operon (Jacob and Monod) in E. coli consists of one regulatory gene (i, the repressor gene) and three structural genes: z (beta-galactosidase, hydrolyses lactose into glucose and galactose), y (permease, increases membrane permeability to allow lactose entry), and a (transacetylase).
- The i gene codes for a repressor protein that, in the absence of lactose (the inducer), binds to the operator region and blocks RNA polymerase from transcribing the structural genes; the operon is therefore "off" by default.
- When lactose is present, it (or its isomer allolactose) acts as an inducer, binding to the repressor and changing its conformation so it can no longer bind the operator. RNA polymerase can then transcribe the z, y, and a genes, producing the enzymes needed for lactose metabolism.
- This is a classic example of a negative inducible control system and shows how gene expression can be switched on or off depending on the physiological requirement of the cell.
Human Genome Project and DNA Fingerprinting
- The Human Genome Project (HGP, 1990-2003), a mega international collaboration, aimed to sequence the entire human genome (about 3 billion base pairs) and identify all the genes it contains.
- Key findings: the human genome contains about 30,000 to 25,000 genes or fewer (far fewer than once predicted); over 90% of the genome consists of non-coding sequences (earlier called "junk DNA"); the average gene size is about 3000 bases, though sizes vary enormously; chromosome 1 has the most genes, and the Y chromosome has the fewest.
- Techniques used included Sequence Annotated Tags (ESTs) and shotgun sequencing of random DNA fragments, followed by alignment using computer programs (bioinformatics).
- DNA fingerprinting exploits the polymorphism (variation) in DNA sequence at specific repetitive regions called satellite DNA / VNTRs (Variable Number of Tandem Repeats), which differ from one individual to another (except identical twins).
- The technique, pioneered by Alec Jeffreys, involves isolating DNA, cutting it with restriction enzymes, separating fragments by gel electrophoresis, and using a radioactive/labelled VNTR probe for hybridisation (Southern blotting) to generate a unique band pattern.
- In India, Dr. Lalji Singh contributed extensively to developing DNA fingerprinting techniques; applications include forensic investigation, paternity testing, and population and genetic diversity studies.
The Search for the Genetic Material
- Griffith's transforming principle (1928): Frederick Griffith worked with Streptococcus pneumoniae. The virulent S (smooth, capsulated) strain killed mice, while the non-virulent R (rough, non-capsulated) strain did not. Heat-killed S bacteria alone were harmless, but a mixture of heat-killed S and live R bacteria killed the mice, and live S bacteria were recovered from them. Griffith concluded that the R strain had been "transformed" by some factor transferred from the dead S strain, though he could not identify the biochemical nature of this transforming principle.
- Avery, MacLeod and McCarty (1933-44) purified biochemicals (proteins, DNA, RNA) from heat-killed S cells to see which one transformed live R cells into S cells. They found that only DNA caused transformation; digestion with DNase abolished transformation, whereas proteases and RNases did not. This indicated that DNA is the hereditary material, though not all biologists were convinced.
- Hershey-Chase experiment (1952): Alfred Hershey and Martha Chase worked with bacteriophages. They grew some phages on a medium containing radioactive phosphorus (32P, which labels DNA) and others on a medium with radioactive sulphur (35S, which labels protein coat). The phages were allowed to infect E. coli; the cultures were then blended and centrifuged. Radioactive DNA (32P) was found inside the bacterial cells, while radioactive protein (35S) remained outside. This unambiguously showed that DNA, not protein, is the genetic material that passes into the host to direct the formation of new phages.
Properties of Genetic Material (DNA versus RNA) and the RNA World
- A molecule that can act as the genetic material must fulfil four criteria: it must be able to replicate; it must be chemically and structurally stable; it must provide the scope for slow mutations (changes) that are required for evolution; and it must be able to express itself in the form of Mendelian characters.
- Both DNA and RNA are able to replicate and mutate. However, RNA is catalytically reactive and therefore unstable. In DNA, the presence of thymine (in place of uracil) and the 2'-deoxyribose sugar (which lacks the reactive 2'-OH group) make DNA chemically less reactive and structurally more stable, so DNA is better suited for the long-term storage of genetic information.
- RNA can also serve as genetic material, as seen in many viruses such as Tobacco Mosaic Virus (TMV), QB bacteriophage and the retroviruses (e.g. HIV).
- RNA world hypothesis: RNA is believed to have been the first genetic material, and it also acted as a catalyst (as ribozymes). Essential life processes such as metabolism, translation and splicing evolved around RNA. Later, the more stable DNA evolved from RNA (with chemical modifications) and took over the role of information storage, while proteins took over most catalytic functions.
🚀 NEET Advanced Edge
Why the lac operon is "negative AND inducible" regulation: It's negative because a repressor protein normally BLOCKS transcription; it's inducible because an external molecule (lactose/allolactose) is needed to RELIEVE that block — contrast with a repressible system where the default is ON and a molecule turns it OFF. Mixing these two terms up is a common exam trap.
Why so few genes are needed: ~20,000-25,000 genes produce a far larger number of proteins through mechanisms like alternative splicing (one gene → multiple mRNA/protein variants) and post-translational modification — this is the key insight resolving the apparent paradox of "fewer genes than expected" from the Human Genome Project.
Worked reasoning: A mutation in the THIRD position of a codon often does NOT change the amino acid produced (e.g. GCU, GCC, GCA, GCG all code for Alanine) — this "wobble"/degeneracy in the third codon position is why many point mutations are silent (no effect on the protein), a frequently tested genetic-code property.