Field guide

Mostly Harmless

A field guide to assumptions: your best friends, or hidden enemies?

An assumption is a trade. You give up generality and get a cleaner argument in return. A good one buys a lot and costs little. A bad one makes the theorem easy to prove and hard to believe.

The labels below are about how reviewers usually read an assumption in a modern ML paper, not about mathematical validity. The same assumption can be fine in a deliberately idealized theorem and unacceptable when the prose stretches the result to realistic deep learning. For each of the 29 assumptions you get what it says, what it buys you, when it is acceptable, a diagnostic you can run on a real model, and the questions where it usually appears.

A diagnostic can refute an assumption, or make it plausible on the region you probed. It never proves it; only a mathematical argument does. Report diagnostics as evidence about your setting, not as verification.

Find an assumption

Filter by verdict or by question

Group 1

Mostly harmless

Standard, widely used, and rarely challenged when stated precisely. Say them clearly and move on.

L-smoothness

Mostly harmless

\(\lVert \nabla f(x) - \nabla f(y) \rVert \le L\lVert x - y \rVert\)

Buys you
The descent lemma, stable step sizes, first-order rates.
Why it's acceptable
Smooth losses and activations provide it on bounded regions. Mostly harmless locally; claim it globally only with care.
Diagnostic
Estimate the Hessian's operator norm \(\max_i \lvert \lambda_i(\nabla^2 f) \rvert\) at checkpoints along the trajectory with Lanczos, looking at both ends of the spectrum: large negative curvature breaks smoothness too. Report a local empirical \(L\) for the region you probed, not "verified global smoothness".
Questions
2, 4

Bounded domain or bounded iterates

Mostly harmless

\(\lVert x - x_0 \rVert \le R\) for every feasible or visited \(x\)

Buys you
Finite Lipschitz and smoothness constants; covering arguments.
Why it's acceptable
Harmless if a projection enforces it, or if you prove it separately.
Diagnostic
Log \(\max_t \lVert x_t - x_0 \rVert\). If projection is part of the algorithm, the assumption holds exactly.
Questions
2, 4

Unbiased stochastic gradients

Mostly harmless

\(\mathbb{E}[g_t \mid x_t] = \nabla f(x_t)\)

Buys you
Cancels the cross terms in SGD analysis.
Why it's acceptable
True for ordinary uniform minibatching.
Diagnostic
At a few checkpoints, compare the mean of many minibatch gradients with the full gradient. Reweighting, importance sampling or compression can introduce bias.
Questions
2

Bounded or normalized inputs

Mostly harmless

\(\lVert x_i \rVert \le B\), often \(\lVert x_i \rVert = 1\)

Buys you
Finite Lipschitz, NTK and margin constants.
Why it's acceptable
Harmless when your preprocessing actually enforces it.
Diagnostic
Plot input norms and document the normalization.
Questions
2, 4

Lipschitz activations

Mostly harmless

\(\lvert \sigma(a) - \sigma(b) \rvert \le L_\sigma \lvert a - b \rvert\)

Buys you
Propagates perturbation and generalization bounds through layers.
Why it's acceptable
ReLU, leaky ReLU, GELU, tanh and sigmoid all satisfy it. Second-derivative assumptions are a separate matter.
Diagnostic
Usually analytic. For learned or custom activations, evaluate the derivative's maximum over the observed pre-activation range.
Questions
4

Random Gaussian initialization

Mostly harmless

\(W^{(0)}_{ij} \overset{\text{iid}}{\sim} \mathcal{N}(0, \sigma_0^2)\) under a stated scaling

Buys you
Random-feature and NTK concentration; symmetry.
Why it's acceptable
Harmless when it is literally what the code does.
Diagnostic
Check the initialization code. The assumption is about time zero, not about the trained weights.
Questions
2

Norm bounds on learned weights

Mostly harmless

\(\lVert W_\ell \rVert_2 \le s_\ell,\ \lVert W_\ell \rVert_F \le F_\ell\)

Buys you
A Lipschitz constant for the network and complexity bounds.
Why it's acceptable
Harmless when the theorem conditions on the norms actually learned; stronger when they are fixed in advance.
Diagnostic
Power iteration for each \(\lVert W_\ell \rVert_2\). Report the actual products, not an abstract constant.
Questions
4
Group 2

Friends, with conditions

Defensible and often necessary, but only if you say where they hold and show some evidence.

Convexity

Justify it

\(f(\lambda x + (1-\lambda)y) \le \lambda f(x) + (1-\lambda) f(y)\)

Buys you
Stationary points are global; Jensen; clean \(\mathcal{O}(1/T)\) rates.
Why it's acceptable
Fine for convex heads, linear models and convex subproblems. A red flag when an end-to-end deep objective is quietly replaced by a convex one.
Diagnostic
Say exactly which subproblem is convex. Chord tests and Hessian eigenvalues can find violations but cannot prove global convexity.
Questions
1, 2, 3

Local strong convexity

Justify it

\(\nabla^2 f(x) \succeq \mu I\) for \(x \in B(x^\star, r)\), possibly on a subspace

Buys you
Local uniqueness, perturbation analysis, local linear convergence.
Why it's acceptable
Much more credible than a global statement, especially with weight decay. Give the radius \(r\).
Diagnostic
Lanczos or Hessian–vector products across checkpoints and nearby random perturbations. Disclose null directions from symmetries.
Questions
1, 2

Polyak–Łojasiewicz inequality

Justify it

\(\tfrac12 \lVert \nabla f(x) \rVert^2 \ge \mu\big(f(x) - f^\star\big)\)

Buys you
Linear convergence in function value without convexity.
Why it's acceptable
Acceptable if proved on the trajectory. Weaker than strong convexity. A red flag if simply postulated globally for a generic deep network.
Diagnostic
Plot \(\lVert \nabla f \rVert^2 / [2(f - f_{\text{best}})]\) along training and under local perturbations; note that \(f_{\text{best}}\) only stands in for \(f^\star\).
Questions
2

Kurdyka–Łojasiewicz property

Justify it

\(\varphi'\big(f(x) - f(x^\star)\big)\operatorname{dist}\big(0, \partial f(x)\big) \ge 1\) near \(x^\star\)

Buys you
Convergence of descent sequences; rates from the KL exponent; distance to minimizer sets.
Why it's acceptable
Holds, with some exponent, for broad classes of losses built from analytic or semialgebraic pieces. The exponent is not automatic, and the rate depends on it: exponent \(\tfrac12\), which implies a square-root error bound near a minimizer set, is a genuine extra assumption. Tie it to your actual architecture and loss.
Diagnostic
Cite the structural result that applies. Empirically you can only fit the implied gradient–gap scaling locally.
Questions
1, 2

Hessian Lipschitzness

Justify it

\(\lVert \nabla^2 f(x) - \nabla^2 f(y) \rVert \le \rho \lVert x - y \rVert\)

Buys you
Taylor remainders, saddle escape, cubic regularization.
Why it's acceptable
Standard in second-order complexity theory; less natural for ReLU networks.
Diagnostic
Estimate \(\lVert H(x+\delta) - H(x) \rVert / \lVert \delta \rVert\) with Hessian–vector products on local perturbations.
Questions
2, 6

Bounded gradient variance

Justify it

\(\mathbb{E}\big[\lVert g_t - \nabla f(x_t) \rVert^2 \mid x_t\big] \le \sigma^2\)

Buys you
The SGD noise floor and \(\mathcal{O}(T^{-1/2})\)-type rates.
Why it's acceptable
Reasonable locally; a uniform global bound can be strong.
Diagnostic
Estimate the minibatch-gradient variance at many checkpoints and report how it changes during training.
Questions
2

I.i.d. data

Justify it

\(Z_1, \dots, Z_n \overset{\text{iid}}{\sim} P\)

Buys you
Product-measure concentration, symmetrization, \(D_{\mathrm{KL}}(P^{\otimes n} \Vert Q^{\otimes n}) = n D_{\mathrm{KL}}(P \Vert Q)\).
Why it's acceptable
Harmless only when the scope is genuinely i.i.d. An important limitation in federated, temporal, clustered, multi-site and domain-shift settings.
Diagnostic
Audit the collection protocol; measure duplicates and group, time or site correlations. Do not claim a statistical test "verified" i.i.d.
Questions
4, 6

Bounded loss

Justify it

\(0 \le \ell(h, z) \le M\)

Buys you
Hoeffding, simple PAC-Bayes and stability concentration.
Why it's acceptable
Harmless for 0-1 or clipped losses. Cross-entropy and squared error are not globally bounded, so justify it otherwise.
Diagnostic
Check the analytic range and the empirical maxima. If you introduce clipping, say how it changes the objective.
Questions
4

Sub-Gaussian noise

Justify it

\(\mathbb{E}\, e^{\lambda X} \le e^{\sigma^2 \lambda^2 / 2}\) for all \(\lambda\), for centered \(X\)

Buys you
Gaussian-like concentration without exact Gaussianity.
Why it's acceptable
Standard and much weaker than Gaussianity.
Diagnostic
Inspect tail plots and QQ plots along random and principal directions; report how sensitive results are to the tails.
Questions
4, 5, 6

Isotropy

Justify it

\(\mathbb{E}X = 0,\ \mathbb{E}[XX^\top] = I\)

Buys you
Removes covariance conditioning; simplifies concentration and spectral analysis.
Why it's acceptable
Reasonable after whitening. Say whether the whitening uses population or sample covariance.
Diagnostic
Compute the empirical covariance spectrum.
Questions
4, 5

Realizability

Justify it

\(\exists h^\star \in \mathcal{H}\) with \(R(h^\star) = 0\), or \(P(Y \mid X)\) lies in the model family

Buys you
Fast rates and no approximation error.
Why it's acceptable
A useful idealization. Interpolating a finite training set does not prove population realizability.
Diagnostic
Report training residuals and misspecification checks on held-out data; call it a theoretical idealization.
Questions
3, 4

Linear separability with margin

Justify it

\(\exists w:\ y_i w^\top x_i \ge \gamma > 0\), often \(\lVert w \rVert = 1\)

Buys you
Max-margin implicit bias and fast classification bounds.
Why it's acceptable
Natural for implicit-bias and interpolation work, and should be stated as part of the result.
Diagnostic
Solve a hard-margin SVM; report the normalized margin and the fraction of violations.
Questions
1, 4

Low-noise (margin) condition

Justify it

\(P\big(\lvert \eta(X) - \tfrac12 \rvert \le t\big) \le C t^\alpha\)

Buys you
Faster excess-risk rates for classification, and sharper surrogate-risk relations.
Why it's acceptable
Encodes a real property of the task distribution, so argue it for your task.
Diagnostic
Fit several probabilistic models and inspect the mass of predicted probabilities near \(1/2\); check sensitivity across estimators.
Questions
3, 4

Restricted strong convexity

Justify it

\(v^\top \nabla^2 f(\theta) v \ge \kappa \lVert v \rVert^2\) for \(v\) in a cone \(\mathcal{C}\)

Buys you
High-dimensional sparse or structured recovery despite a globally singular Hessian.
Why it's acceptable
Acceptable when the cone is explicit.
Diagnostic
Estimate minimum Rayleigh quotients over sampled sparse or tangent directions; for linear models, inspect restricted covariance spectra.
Questions
1, 4, 5

Low-rank structure

Justify it

\(\operatorname{rank}(M^\star) \le r \ll \min(m, n)\)

Buys you
Reduced sample complexity; nuclear-norm and factorized parameterizations.
Why it's acceptable
Often scientifically interpretable, but approximate low rank should be stated as such.
Diagnostic
Plot the singular-value spectrum, the effective rank and reconstruction error against \(r\).
Questions
4, 5

Incoherence

Justify it

\(\mu(U) = \tfrac{d}{r} \max_i \lVert U^\top e_i \rVert^2 \le \mu_0\)

Buys you
Prevents low-rank signal from concentrating on a few coordinates; enables matrix completion.
Why it's acceptable
Acceptable with justification; it excludes genuinely spiky factors.
Diagnostic
Compute the empirical coherence of the singular vectors and compare with the theorem's threshold.
Questions
5

Overparameterization

Justify it

width \(m \ge \operatorname{poly}(n, L, 1/\lambda_0, \log(1/\delta), \dots)\)

Buys you
A concentrated Gram or NTK matrix, favorable local geometry, interpolation.
Why it's acceptable
Acceptable, but reviewer-sensitive when the threshold is orders of magnitude beyond the experiments.
Diagnostic
Plug your actual \(n\), depth, \(\lambda_0\) and \(\delta\) into the threshold and report the gap; run width sweeps.
Questions
2, 4

Identifiability and non-singular information

Justify it

\(P_\theta = P_{\theta'} \Rightarrow \theta = \theta'\) modulo a stated group \(G\); locally \(I(\theta^\star) \succ 0\)

Buys you
Unique recovery, asymptotic normality and efficiency.
Why it's acceptable
Acceptable once the symmetries are named. A red flag if they are ignored.
Diagnostic
Compute the empirical Fisher or Jacobian spectrum after removing known permutation, scale or gauge directions; check recovery across seeds in simulation.
Questions
4, 5
Group 3

Hidden enemies

These make proofs easy and results weak when they drive a claim about realistic deep learning. They are not always wrong, but they need a strong scope statement.

Global strong convexity of a deep network

Red flag

\(f(y) \ge f(x) + \langle \nabla f(x), y - x \rangle + \tfrac{\mu}{2}\lVert y - x \rVert^2\) for all \(x, y\)

Buys you
A unique optimum, \(\mathcal{O}((1 - \mu/L)^t)\) rates and \(\varepsilon/\mu\) sensitivity.
When it's a problem
Parameter symmetries alone rule it out for full networks. Mostly harmless for regularized convex models.
Diagnostic
State it on a convex subproblem or a local quotient space instead, and estimate the smallest Hessian eigenvalues near the solution after removing symmetry directions.
Questions
1, 2

Globally bounded gradients on \(\mathbb{R}^d\)

Red flag

\(\lVert \nabla f(x) \rVert \le G\) for all \(x\)

Buys you
Lipschitzness and bounded update sensitivity.
When it's a problem
Incompatible with global strong convexity, since a strongly convex function has unbounded gradients. Fine on a bounded region.
Diagnostic
Measure per-example and full-batch gradient norms across training and state the domain explicitly.
Questions
2, 4

Exact Gaussian inputs for claims about real data

Red flag

\(X \sim \mathcal{N}(\mu, \Sigma)\)

Buys you
Rotation invariance, Stein identities, exact random-matrix calculations.
When it's a problem
Natural in teacher–student and other stylized theory. A red flag only when the result is presented as a claim about arbitrary real data.
Diagnostic
State the theorem as a Gaussian-model result. Normality tests and random-projection QQ plots can diagnose severe violations, not certify Gaussianity.
Questions
4, 5, 6

The lazy (NTK) regime as an explanation of feature learning

Red flag

\(\lambda_{\min}(K_0) \ge \lambda_0 > 0\) and \(\sup_t \lVert K_t - K_0 \rVert_2 \le c\lambda_0,\ c < 1\)

Buys you
Linearized dynamics and geometric decay of the residual.
When it's a problem
Correct for a lazy-training theorem. In the infinite-width limit the kernel stays fixed during training, so it cannot explain features that change.
Diagnostic
Compute the empirical NTK on a manageable subset: its smallest eigenvalue, condition number and relative drift during training.
Questions
2, 4
What reviewers say

Recurring objections

Habits

Keeping your assumptions friendly

  1. Start simple, then relax. State a clean didactic theorem that exposes the mechanism, then a relaxation, or say plainly that none is available. In the CLTR paper, equal class sizes within Head and Tail (Assumption 3.2) give the clean \(\mathcal{O}(1/\sqrt{\mathrm{IF}})\) form; Assumptions 3.4 and 3.6 then relax convexity and the equal-size condition. Global convexity was exactly what a reviewer flagged as making the theory inapplicable to real models, and the relaxation is what answers it.
  2. Make assumptions local when the proof only needs locality. "For every \(x\) in the radius-\(r\) neighborhood containing the iterates" is far better science than "for all \(x \in \mathbb{R}^d\)".
  3. Measure the constants. Hessian eigenvalues, NTK gaps, spectral norms, margins, gradient variance and effective rank can all be estimated instead of hidden in \(C\). The CLTR paper estimates the threshold in its guarantee on a trained ResNet-18 (Appendix G.5) instead of only asserting that it is met.
  4. Say what the theorem does not establish. Lazy-training theory does not establish feature learning; a Gaussian teacher–student theorem does not prove ImageNet realism; a stationarity rate is not global optimization; an i.i.d. bound is not a domain-shift guarantee.

Good theory papers earn credibility not by avoiding assumptions, but by stating their scope, exposing the constants, and separating the idealized mechanism from the empirical regime it is meant to illuminate.

← Back to the guide

Found an error, or have a better example or a question this guide should cover? Email m.molahasani.m@gmail.com. Corrections and contributions are welcome. Last updated October 5, 2026.