Convexity
Justify it
\(f(\lambda x + (1-\lambda)y) \le \lambda f(x) + (1-\lambda) f(y)\)
- Buys you
- Stationary points are global; Jensen; clean \(\mathcal{O}(1/T)\) rates.
- Why it's acceptable
-
Fine for convex heads, linear models and convex subproblems. A red flag when an end-to-end deep objective is
quietly replaced by a convex one.
- Diagnostic
-
Say exactly which subproblem is convex. Chord tests and Hessian eigenvalues can find violations but cannot
prove global convexity.
- Questions
- 1, 2, 3
Local strong convexity
Justify it
\(\nabla^2 f(x) \succeq \mu I\) for \(x \in B(x^\star, r)\), possibly on a subspace
- Buys you
- Local uniqueness, perturbation analysis, local linear convergence.
- Why it's acceptable
- Much more credible than a global statement, especially with weight decay. Give the radius \(r\).
- Diagnostic
-
Lanczos or Hessian–vector products across checkpoints and nearby random perturbations. Disclose null
directions from symmetries.
- Questions
- 1, 2
Polyak–Łojasiewicz inequality
Justify it
\(\tfrac12 \lVert \nabla f(x) \rVert^2 \ge \mu\big(f(x) - f^\star\big)\)
- Buys you
- Linear convergence in function value without convexity.
- Why it's acceptable
-
Acceptable if proved on the trajectory. Weaker than strong convexity. A red flag if simply postulated globally
for a generic deep network.
- Diagnostic
-
Plot \(\lVert \nabla f \rVert^2 / [2(f - f_{\text{best}})]\) along training and under local perturbations;
note that \(f_{\text{best}}\) only stands in for \(f^\star\).
- Questions
- 2
Kurdyka–Łojasiewicz property
Justify it
\(\varphi'\big(f(x) - f(x^\star)\big)\operatorname{dist}\big(0, \partial f(x)\big) \ge 1\) near \(x^\star\)
- Buys you
- Convergence of descent sequences; rates from the KL exponent; distance to minimizer sets.
- Why it's acceptable
-
Holds, with some exponent, for broad classes of losses built from analytic or semialgebraic pieces. The
exponent is not automatic, and the rate depends on it: exponent \(\tfrac12\), which implies a square-root
error bound near a minimizer set, is a genuine extra assumption. Tie it to your actual architecture and loss.
- Diagnostic
-
Cite the structural result that applies. Empirically you can only fit the implied gradient–gap scaling
locally.
- Questions
- 1, 2
Hessian Lipschitzness
Justify it
\(\lVert \nabla^2 f(x) - \nabla^2 f(y) \rVert \le \rho \lVert x - y \rVert\)
- Buys you
- Taylor remainders, saddle escape, cubic regularization.
- Why it's acceptable
- Standard in second-order complexity theory; less natural for ReLU networks.
- Diagnostic
-
Estimate \(\lVert H(x+\delta) - H(x) \rVert / \lVert \delta \rVert\) with Hessian–vector products on
local perturbations.
- Questions
- 2, 6
Bounded gradient variance
Justify it
\(\mathbb{E}\big[\lVert g_t - \nabla f(x_t) \rVert^2 \mid x_t\big] \le \sigma^2\)
- Buys you
- The SGD noise floor and \(\mathcal{O}(T^{-1/2})\)-type rates.
- Why it's acceptable
- Reasonable locally; a uniform global bound can be strong.
- Diagnostic
- Estimate the minibatch-gradient variance at many checkpoints and report how it changes during training.
- Questions
- 2
I.i.d. data
Justify it
\(Z_1, \dots, Z_n \overset{\text{iid}}{\sim} P\)
- Buys you
-
Product-measure concentration, symmetrization, \(D_{\mathrm{KL}}(P^{\otimes n} \Vert Q^{\otimes n}) = n
D_{\mathrm{KL}}(P \Vert Q)\).
- Why it's acceptable
-
Harmless only when the scope is genuinely i.i.d. An important limitation in federated, temporal, clustered,
multi-site and domain-shift settings.
- Diagnostic
-
Audit the collection protocol; measure duplicates and group, time or site correlations. Do not claim a
statistical test "verified" i.i.d.
- Questions
- 4, 6
Bounded loss
Justify it
\(0 \le \ell(h, z) \le M\)
- Buys you
- Hoeffding, simple PAC-Bayes and stability concentration.
- Why it's acceptable
-
Harmless for 0-1 or clipped losses. Cross-entropy and squared error are not globally bounded, so justify it
otherwise.
- Diagnostic
-
Check the analytic range and the empirical maxima. If you introduce clipping, say how it changes the
objective.
- Questions
- 4
Sub-Gaussian noise
Justify it
\(\mathbb{E}\, e^{\lambda X} \le e^{\sigma^2 \lambda^2 / 2}\) for all \(\lambda\), for centered \(X\)
- Buys you
- Gaussian-like concentration without exact Gaussianity.
- Why it's acceptable
- Standard and much weaker than Gaussianity.
- Diagnostic
-
Inspect tail plots and QQ plots along random and principal directions; report how sensitive results are to the
tails.
- Questions
- 4, 5, 6
Isotropy
Justify it
\(\mathbb{E}X = 0,\ \mathbb{E}[XX^\top] = I\)
- Buys you
- Removes covariance conditioning; simplifies concentration and spectral analysis.
- Why it's acceptable
- Reasonable after whitening. Say whether the whitening uses population or sample covariance.
- Diagnostic
- Compute the empirical covariance spectrum.
- Questions
- 4, 5
Realizability
Justify it
\(\exists h^\star \in \mathcal{H}\) with \(R(h^\star) = 0\), or \(P(Y \mid X)\) lies in the model family
- Buys you
- Fast rates and no approximation error.
- Why it's acceptable
- A useful idealization. Interpolating a finite training set does not prove population realizability.
- Diagnostic
-
Report training residuals and misspecification checks on held-out data; call it a theoretical idealization.
- Questions
- 3, 4
Linear separability with margin
Justify it
\(\exists w:\ y_i w^\top x_i \ge \gamma > 0\), often \(\lVert w \rVert = 1\)
- Buys you
- Max-margin implicit bias and fast classification bounds.
- Why it's acceptable
- Natural for implicit-bias and interpolation work, and should be stated as part of the result.
- Diagnostic
- Solve a hard-margin SVM; report the normalized margin and the fraction of violations.
- Questions
- 1, 4
Low-noise (margin) condition
Justify it
\(P\big(\lvert \eta(X) - \tfrac12 \rvert \le t\big) \le C t^\alpha\)
- Buys you
- Faster excess-risk rates for classification, and sharper surrogate-risk relations.
- Why it's acceptable
- Encodes a real property of the task distribution, so argue it for your task.
- Diagnostic
-
Fit several probabilistic models and inspect the mass of predicted probabilities near \(1/2\); check
sensitivity across estimators.
- Questions
- 3, 4
Restricted strong convexity
Justify it
\(v^\top \nabla^2 f(\theta) v \ge \kappa \lVert v \rVert^2\) for \(v\) in a cone \(\mathcal{C}\)
- Buys you
- High-dimensional sparse or structured recovery despite a globally singular Hessian.
- Why it's acceptable
- Acceptable when the cone is explicit.
- Diagnostic
-
Estimate minimum Rayleigh quotients over sampled sparse or tangent directions; for linear models, inspect
restricted covariance spectra.
- Questions
- 1, 4, 5
Low-rank structure
Justify it
\(\operatorname{rank}(M^\star) \le r \ll \min(m, n)\)
- Buys you
- Reduced sample complexity; nuclear-norm and factorized parameterizations.
- Why it's acceptable
- Often scientifically interpretable, but approximate low rank should be stated as such.
- Diagnostic
- Plot the singular-value spectrum, the effective rank and reconstruction error against \(r\).
- Questions
- 4, 5
Incoherence
Justify it
\(\mu(U) = \tfrac{d}{r} \max_i \lVert U^\top e_i \rVert^2 \le \mu_0\)
- Buys you
- Prevents low-rank signal from concentrating on a few coordinates; enables matrix completion.
- Why it's acceptable
- Acceptable with justification; it excludes genuinely spiky factors.
- Diagnostic
- Compute the empirical coherence of the singular vectors and compare with the theorem's threshold.
- Questions
- 5
Overparameterization
Justify it
width \(m \ge \operatorname{poly}(n, L, 1/\lambda_0, \log(1/\delta), \dots)\)
- Buys you
- A concentrated Gram or NTK matrix, favorable local geometry, interpolation.
- Why it's acceptable
- Acceptable, but reviewer-sensitive when the threshold is orders of magnitude beyond the experiments.
- Diagnostic
-
Plug your actual \(n\), depth, \(\lambda_0\) and \(\delta\) into the threshold and report the gap; run width
sweeps.
- Questions
- 2, 4
Identifiability and non-singular information
Justify it
\(P_\theta = P_{\theta'} \Rightarrow \theta = \theta'\) modulo a stated group \(G\); locally \(I(\theta^\star)
\succ 0\)
- Buys you
- Unique recovery, asymptotic normality and efficiency.
- Why it's acceptable
- Acceptable once the symmetries are named. A red flag if they are ignored.
- Diagnostic
-
Compute the empirical Fisher or Jacobian spectrum after removing known permutation, scale or gauge directions;
check recovery across seeds in simulation.
- Questions
- 4, 5