Guide
The Hitchhiker's Guide to Theoretical Machine Learning
Where to begin
Most ML researchers can follow a proof. Far fewer know where to start when they have to write one. The hard part is usually not the algebra. It is knowing which kind of argument fits the claim you want to make.
This guide is organized around that first decision. Find the question your theorem is trying to answer, and each page tells you which tools usually answer it, what each tool costs you in assumptions, and what real proofs built from them look like.
Don't Panic.
What are you trying to prove?
Pick the question closest to your claim. The tree narrows it down and shows the tools you will probably need.
First-order and KKT conditions · strong-convexity sensitivity · implicit function theorem · implicit-bias analysis
Costs: Differentiability; curvature (or a KL property) near the solution
Explore →Descent lemma · PL inequality · telescoping potentials · SGD recursions · saddle escape
Costs: Smoothness on the trajectory; step-size conditions
Explore →Pointwise domination · ψ-transform calibration · proper-scoring identities · Pinsker · Jensen and ELBO
Costs: Conditional-risk analysis; the right function class
Explore →Concentration · symmetrization and Rademacher · uniform stability · PAC-Bayes · mutual information
Costs: i.i.d. data; bounded or sub-Gaussian loss
Explore →SVD · Eckart–Young · Weyl · Davis–Kahan · Courant–Fischer · identifiability
Costs: An eigengap; an explicit symmetry group
Explore →Le Cam's two-point method · Fano · packings · KL tensorization · Yao's principle · zero-chain functions
Costs: A precisely stated model, oracle and loss
Explore →The first move for each question
- Comparing two endpoints? Subtract their optimality equations.
- Need an optimization rate? Find a quantity that decreases by a definite amount at every step.
- Relating two objectives? Condition on the input and compare the conditional risks.
- Finite-data performance? First decide between a bound that is uniform over the model class and one that is specific to your algorithm. Only then pick a concentration tool.
- Describing what features look like? Find the matrix or equivalence relation whose structure is the claim.
- Claiming nobody can do better? Write down the minimax quantifiers, then build instances that are far apart in your loss but hard to tell apart from data.
Robustness, shift, privacy and other modifiers
Some guarantees are not a seventh question. They change one of the six. A robustness theorem might ask whether a surrogate certifies the adversarial loss (Question 3) or whether a predictor generalizes over perturbations (Question 4). Tag your theorem with one primary question and as many modifiers as apply.
| Modifier | Usually lands in |
|---|---|
| Robustness and adversarial guarantees | Question 3 (certify the robust loss), 4 (generalize over perturbations), 6 (what robustness costs) |
| Distribution shift | Question 4 with the i.i.d. assumption replaced, or Question 5 when the claim is about what stays invariant |
| Privacy | Question 4 (utility under privacy) and Question 6 (how privacy changes the best possible rate) |
| Causal claims | Question 5 for identifiability, then Question 4 for estimation |
| Scaling laws | Question 2, 4 or 6, depending on whether the quantity is optimization error, statistical risk or a fundamental limit |
Assumptions: best friends or hidden enemies?
Every theorem rests on assumptions, and they are usually what reviewers argue about. Some cost you nothing. Some are fine if you justify them. Some quietly make your result say much less than it seems to.
One real paper, from observation to theorem
The chapters split the work into questions and tools. The case study puts it back together: one result followed from the experiment that suggested it, through the first assumption and the objection to it, to the relaxed theorem and the second result it made possible.
Tool index
The same map read the other way: each tool, the questions where it usually appears, and where to learn it.
Reading a question page
- The question and the symptoms that tell you it is yours.
- Sub-cases: the narrower versions you will usually face, and where each one starts.
- The recipe: the proof skeleton most answers follow.
- A complete proof, step by step with commentary, followed by a tempting approach that fails and what changes when you relax an assumption.
- An exercise with a hint and a hidden solution.
- The toolkit: each tool's statement, what it gives you, what it costs and when it is the wrong tool.
- Worked examples from published papers, including my own where they fit.
- Traps reviewers look for, and resources to go deeper.
Theorem numbers are given only where they were checked against the paper; elsewhere the guide refers to a paper's main result. Several inequalities have more than one common normalization. Each page states one convention; when you import a result, keep the source paper's convention throughout your proof.
Found an error, or have a better example or a question this guide should cover? Email m.molahasani.m@gmail.com. Corrections and contributions are welcome. Last updated October 5, 2026.