Physical Chemistry · Part 9 of 9

Computation, Spectroscopy & the Quantum Toolkit

Quantum Chemistry, Part 9 · 15 sections · about 25,648 words · CSIR-NET Chemical Sciences, GATE Chemistry & IIT-JAM

The last part connects the machinery to the two places a chemist actually meets it: the computer and the spectrometer. It explains what a basis set is and what the acronyms on a paper's methods line mean, derives the transition-moment integral that every selection rule in spectroscopy comes from, shows how symmetry decides whether an integral vanishes without computing it, and closes with the master index to all nine parts. Two layers on every section: a slow, hand-held beginner path and a research-grade advanced/reference path.

The 15 sections in Part 9

  • 1Basis sets — the vocabulary the molecule is described in Free below
  • 2Past Hartree–Fock — the correlation energy and how to get it back
  • 3Density functional theory
  • 4Semi-empirical methods, geometry optimisation, and reading an output file
  • 5The transition dipole moment — where every selection rule comes from
  • 6Rotational and vibrational spectra — the rules derived
  • 7Electronic spectra and the Franck–Condon principle
  • 8Raman, magnetic resonance and photoelectron spectroscopy
  • 9The master formula sheet
  • 10Every exactly solvable system, side by side
  • 11Units, constants and conversions
  • 12Which method for which problem
  • 13The top exam traps, Parts 1–9
  • 14Section-by-section index
  • 15The argument, from Part 1 to Part 9

Basis sets — the vocabulary the molecule is described in

Free extract

Section J.1 of Part 9, reproduced in full from the book — figures and all. No sign-in, no paywall on this section.

Beginner layer

Start here. A basis set is a vocabulary, and the whole subject follows from one algebraic accident about Gaussians.

Part 8 built molecular orbitals as linear combinations of atomic orbitals: ψ = Σi ciφi. The variational method of Part 6 then finds the best coefficients ci. But it finds the best coefficients only for the functions φi you gave it. If the right answer cannot be written as a combination of your φi, no amount of optimisation will produce it.

Basis set: the fixed set of one-electron functions from which every molecular orbital in a calculation is built. It is the vocabulary of the calculation. The variational principle guarantees the best sentence that can be written in that vocabulary — and nothing about whether the vocabulary was adequate.

So the first question is which functions to use. There are two natural candidates and they differ in a single exponent.

Slater-type orbital (STO):   φ ∝ rn−1 e−ζr Ylm(θ,φ)
Gaussian-type orbital (GTO):   g ∝ xaybzc e−αr2

The Slater function is the right shape. It is the exact form of the hydrogen 1s orbital of Part 5, it has the correct cusp at the nucleus, and it decays exponentially, which is what the true wavefunction of a bound electron does far from the nuclei. The Gaussian is the wrong shape in both places: it is flat at the nucleus instead of pointed, and it dies far too fast in the tail.

Why chemistry is done with the wrong-shaped function012340.00.20.40.6φ(r) / a₀−3/2Slater 1s, e−ζr — the right shape1 Gaussian   overlap 0.97702 Gaussians   overlap 0.99843 Gaussians   overlap 0.9996The cusp — where every Gaussian fails00.20.40.350.450.55STO: dφ/dr = −ζφ at r = 0 — a sharp point.Every Gaussian is flat there. No sum fixes it.The tail — log scale471010⁻²10⁻⁵10⁻⁸e−ζr is a straight line here; e−αr² bends away.At r = 10a₀ the 3G fit is already 29× too small.r / a₀The Gaussian curves are least-squares fits computed at build time bymaximising the overlap with a Slater function over an even-temperedexponent sequence. They are our fits, not the published STO-nGcoefficients. By overlap: 1G captures 97.70%, 2G 99.84%, 3G 99.958%.
A Gaussian is the wrong shape in both directions, and it won anyway. The Slater function has a cusp at the nucleus, as the Coulomb singularity demands, and a slowly decaying tail. A Gaussian is flat at the origin and dies far too fast outside; neither defect is repairable by any finite sum. What makes Gaussians unbeatable is one algebraic fact: the product of two Gaussians on different atoms is a single Gaussian in between, turning the four-centre two-electron integrals into closed-form expressions. So we fix the shape by brute force: contract several Gaussians into one fixed combination that imitates a Slater function.

And yet essentially every molecular calculation done in the last fifty years uses Gaussians. The reason is one theorem.

The Gaussian product theorem: the product of two Gaussian functions centred on different points is a single Gaussian centred on a point between them. Symbolically, e−α|r−A|² × e−β|r−B|² = K e−(α+β)|r−P|² with P the weighted midpoint of A and B and K a constant.

That collapses a four-centre two-electron integral — four different atoms, the computational bottleneck of every ab initio calculation — into a two-centre integral with a closed-form answer. With Slater functions those integrals must be evaluated numerically, one at a time, and there are of order N4 of them. The Gaussian is a worse function that makes a vastly better program. Boys saw this in 1950 and the field has never looked back.

The fix for the wrong shape is brute force. Take several Gaussians of different widths, fix their relative weights once and for all, and use the whole combination as a single basis function.

Contracted Gaussian function (CGF): a fixed linear combination of primitive Gaussians, χ = Σk dkgk, in which the contraction coefficients dk and the exponents αk are determined in advance and held fixed during the molecular calculation. Only the coefficients ci multiplying whole contracted functions are varied.
Number of primitivesExponents α (our fit)Coefficients (our fit)Overlap with the Slater function
10.30221.00000.97698
20.1441, 0.78540.6502, 0.45840.99838
30.1126, 0.4560, 1.84680.4904, 0.4954, 0.14490.99958
Computed at build time by maximising the overlap of the normalised contraction with a ζ = 1 Slater 1s function over an even-tempered exponent sequence. These are our own numbers, generated by the build script — they are not the published STO-3G parameters and must not be quoted as such. The trend is the teaching point: three Gaussians already reproduce the Slater function to better than 99.96%.

The four ways a basis set gets bigger

Once contraction is settled, every remaining decision is about how many functions to use and of what kind. There are four independent axes and each fixes a specific physical deficiency.

  1. Minimal basis. One contracted function per occupied atomic orbital. Carbon gets 1s, 2s, 2px, 2py, 2pz — five functions. Hydrogen gets one. STO-3G is the classic. Defect: the size of every orbital is frozen, so an atom cannot become more compact when it gains positive charge or more diffuse when it gains negative charge.
  2. Double zeta / split valence. Two functions per orbital, one tight and one loose. Their ratio is now variational, so the orbital can breathe. Full double zeta doubles the core as well, which is wasteful because core orbitals barely change on bonding; split-valence sets double only the valence and leave the core minimal. 6-31G and 3-21G are split-valence.
  3. Polarisation functions. Functions of one higher angular momentum than the atom needs: d functions on carbon, nitrogen, oxygen; p functions on hydrogen. Defect they fix: an s orbital on hydrogen is perfectly spherical and can only sit on the nucleus. Mixing in a little p lets the density shift off-centre, which is what actually happens when a bond forms. Without polarisation functions, bond angles in strained rings, hypervalent geometries and the entire barrier to inversion in NH3 come out wrong.
  4. Diffuse functions. Extra functions with very small exponents, so they extend far from the nucleus. Marked by + in Pople notation and aug- in Dunning notation. Essential for anions (where the extra electron is loosely held), lone pairs, excited and Rydberg states, and any weak intermolecular interaction. Omitting them from an anion calculation is one of the commonest errors in the literature.
How to read a basis-set nameAnatomy of a Pople basis-set name6-311++G(2df,2p)12345616 primitive Gaussianscontracted into ONE core function2valence split THREE ways3 + 1 + 1 primitives → three functions3diffuse functionsfirst + on heavy atoms, second + on H4‘G’ = Gaussianthe family marker5polarisation on heavy atomstwo sets of d, one set of f6polarisation on hydrogentwo sets of pHow a basis set grows — each step, and what it buys. The first bar is the baseline; the last is a family, not an axis.MINIMALone function per occupied AOSTO-3Gcannot change the size of an orbital at allDOUBLE-ZETA VALENCEtwo functions per valence AO6-31Gan orbital can now contract or expand+ POLARISATIONadd l+1 functions: d on C/N/O, p on H6-31G(d,p)an orbital can now BEND — bond angles, strained rings+ DIFFUSEadd small-exponent functions6-31+G(d)anions, lone pairs, Rydberg states, weak interactionsCORRELATION-CONSISTENTshells added in balanced setscc-pVXZthe energy converges smoothly and can be extrapolatedThe bar length is a visual guide to relative size only. The actual function counts for water are computed in the table below the figure.
The name is a specification, not a brand. Read a Pople name left to right: the number before the hyphen is how many primitives are contracted into each core function; the digits after it say how the valence shell is split; plus signs add diffuse functions (one = heavy atoms, two = hydrogens too); the bracket lists polarisation functions, heavy atoms before the comma. The older asterisks mean the same: 6-31G* is 6-31G(d), 6-31G** is 6-31G(d,p). Dunning's cc-pVXZ sets matter not because they are larger but because they form a sequence that can be extrapolated to the complete-basis limit.

Now the arithmetic. Here is the same molecule — water, ten electrons — in eight standard basis sets, with the number of contracted basis functions counted by the build script from the published contraction patterns.

Basis setCharacterBasis functions for H2ONote
STO-3Gminimal7one function per occupied atomic orbital, nothing more
3-21Gsplit valence (double-zeta valence)13core single, valence split into two
6-31Gsplit valence (double-zeta valence)13same size as 3-21G, better primitives
6-31G(d)+ d polarisation on heavy atoms196 Cartesian d functions on O
6-31G(d,p)+ p polarisation on H25the standard workhorse
6-311++G(d,p)triple-split valence + diffuse on all atoms36core 1 + valence 3 + diffuse 1 for each of s and p; 5 pure d — the 6-311 family default, unlike 6-31G's 6 Cartesian d
cc-pVDZcorrelation-consistent double zeta245 pure d functions, designed to recover correlation energy
cc-pVTZcorrelation-consistent triple zeta58the first basis worth extrapolating from
Counted at build time by summing shell sizes (s = 1, p = 3, pure d = 5, Cartesian d = 6, pure f = 7) over the published contraction pattern for each atom, each row in the convention its own basis family uses: the 6-31G sets take 6 Cartesian d, the 6-311G sets and the Dunning cc-pVXZ sets take 5 pure d. Count 6-311++G(d,p) with 6d instead and you get 37, which is not the number the program will print. Note the jump from 7 to 25 functions in going from a minimal basis to the standard workhorse, and to 58 for a triple-zeta correlation-consistent set — and remember that the number of two-electron integrals goes as the fourth power of that count.
Advanced / reference layer

Basis-set superposition error, the counterpoise correction, and the complete-basis-set limit.

Basis-set superposition error (BSSE): in a calculation on a complex A⋅B, the basis functions belonging to B are physically present near A and improve A's description, and vice versa. Each monomer is therefore described better inside the complex than alone, which artificially stabilises the complex. The error is largest for small basis sets and for weak interactions — precisely where interaction energies are being computed.

The standard repair is the counterpoise correction of Boys and Bernardi: compute each monomer in the full basis of the complex, with the other monomer's nuclei removed but its basis functions ('ghost functions') retained. Then

ΔECP = EAB(AB basis) − EA(AB basis) − EB(AB basis)

Counterpoise usually overcorrects, so a common convention is to quote both the corrected and uncorrected values and treat their difference as an error bar. BSSE vanishes only in the complete-basis limit, which suggests the other repair: extrapolate.

Complete-basis-set extrapolation. The correlation-consistent sets were designed so that the correlation energy converges smoothly with the cardinal number X (D = 2, T = 3, Q = 4, 5 = 5). The Hartree–Fock energy converges roughly exponentially in X; the correlation energy converges as X−3, a consequence of how slowly a product basis describes the electron–electron cusp. The standard two-point formula is

Ecorr(X) = Ecorr(∞) + A X−3  ⇒  Ecorr(∞) = [X³Ecorr(X) − Y³Ecorr(Y)] / (X³ − Y³)
Symptom in the outputLikely basis-set causeRepair
Anion binding energy far too small, or the anion is unboundno diffuse functionsadd + (or aug-)
Bond angles at strained or hypervalent centres badly wrongno polarisation functionsadd (d) — and (d,p) if H is involved in bonding
Interaction energy of a dimer much too largeBSSE with a small basiscounterpoise correction, or a much larger basis
Two programs disagree in the same nominal basis6d versus 5d Cartesian/pure conventionstate the convention explicitly
SCF fails with a ‘linear dependence’ warningover-complete diffuse setcanonical orthogonalisation, or drop functions
Energies of a heavy-element compound are absurdno ECP, no relativistic treatmentuse an ECP-matched basis
The diagnostic table for J.1. Most basis-set failures announce themselves; the dangerous ones are BSSE and the missing diffuse function, which produce plausible numbers that happen to be wrong.

Decode 6-311++G(2df,2p) completely, and count the basis functions for CH3OH Medium

Step 1 — the core. The 6 before the hyphen: each core orbital is a single contracted function built from six primitive Gaussians. Carbon and oxygen each get one core function (their 1s).
Step 2 — the valence split. 311 after the hyphen: the valence is split three ways, into contracted functions of 3, 1 and 1 primitives. So each valence orbital (2s and each 2p) is represented by three functions. This is a triple-zeta valence basis.
Step 3 — the diffuse functions. ++: the first plus adds one diffuse s and one diffuse p shell to each heavy atom; the second plus adds a diffuse s to each hydrogen.
Step 4 — the polarisation functions. (2df,2p): before the comma, heavy atoms get two sets of d functions and one set of f. After the comma, hydrogens get two sets of p.
Step 5 — count for carbon (and identically for oxygen). Core 1s: 1 function. Valence: 4 orbitals (2s, 2p×3) × 3 splittings = 12. Diffuse: 1 s + 3 p = 4. Polarisation: 2 sets of d and 1 set of f. Here the convention decides the answer. The 6-311 family is a 5d/7f (pure) basis, which is what Gaussian and every other program use for it by default: 2 × 5 + 7 = 17. Total = 1 + 12 + 4 + 17 = 34 per heavy atom.
Step 6 — count for hydrogen. Valence 1s split three ways = 3. Diffuse s = 1. Two sets of p = 6. Total = 10 per hydrogen — the same in either convention, because p functions have no pure/Cartesian split.
Step 7 — assemble CH3OH. Two heavy atoms (C, O) and four hydrogens: 2 × 34 + 4 × 10 = 108 contracted basis functions. The number of two-electron integrals scales as N⁴/8 ≈ 1.7 × 107. That is a small calculation by modern standards and an impossible one in 1970.
Step 8 — the answer you would get in the other convention, and why it matters. Forced to 6d/10f Cartesian, the same set gives 2 × 6 + 10 = 22 polarisation functions, 1 + 12 + 4 + 22 = 39 per heavy atom and 118 for the molecule. Ten extra functions on two atoms is not a rounding error: it changes the energy, and two people quoting ‘6-311++G(2df,2p)’ without saying 5d or 6d are not doing the same calculation. Say which.

⚠ Common mistakes & exam traps

  • Believing that a bigger basis set always gives a better answer. It gives a better answer to the equations you are solving. HF/cc-pV5Z converges beautifully to the Hartree–Fock limit, which is not the experimental answer, and never will be: the missing correlation energy is a fixed deficit no basis can repair.
  • Confusing the number of primitives with the number of basis functions. STO-3G uses three primitives per function, but the variational problem has one coefficient per contracted function. Water in STO-3G has 21 primitives and 7 basis functions, and it is the 7 that sets the cost of the diagonalisation.
  • Running an anion without diffuse functions. The extra electron in F or an enolate sits well outside the neutral density. Without small-exponent functions the basis cannot describe it, and electron affinities come out far too small — sometimes negative.
  • Quoting an interaction energy without saying whether it is counterpoise corrected. For a hydrogen-bonded dimer in a modest basis, BSSE can be a substantial fraction of the interaction energy itself.
  • Assuming ‘double zeta’ and ‘split valence’ are synonyms. A true double-zeta basis doubles the core as well. 6-31G doubles only the valence and is a split-valence basis; the distinction is a favourite one-mark question.
  • Forgetting that basis functions live on atoms, not on molecules. The basis moves with the nuclei during a geometry optimisation, which is why the gradient of the energy contains Pulay forces — terms from the derivative of the basis functions themselves, not just of the Hamiltonian.

Read the rest of Part 9

The remaining 14 sections of this part — Past Hartree–Fock — the correlation energy and how to get it back, Density functional theory, Semi-empirical methods, geometry optimisation, and reading an output file, The transition… — and all nine parts of Quantum Chemistry are part of ChemVidya Full Access, along with the other books, 55 Study Notes and 6,000+ practice questions.

See plans Open in the app

Continue through Quantum Chemistry

Related Physical Study Notes