Pointwise domination
For a binary margin \(m = y f(x)\).
- Gives
- An immediate risk upper bound.
- Costs
- A true pointwise inequality.
- Wrong tool when
- What matters is excess risk or calibration, not the absolute risk.
Question 3
Your algorithm optimizes a convenient loss or divergence A. Does a small A provably mean a small error in the metric B you actually care about?
| Case | Usually start with |
|---|---|
| Is a pointwise upper bound enough? | Direct domination, for example the 0-1 loss is at most the hinge loss. |
| Need Bayes consistency, not just an upper bound? | Conditional-risk analysis and the calibration (ψ-transform) inequality. |
| A divergence stands in for a distributional discrepancy? | Pinsker, data processing, variational representations, proper-scoring identities. |
| A self-supervised objective should imply downstream performance? | Latent-variable or augmentation assumptions, an explicit downstream classifier, and a transfer bound. |
| A variational lower or upper objective? | Jensen, Fenchel duality, Donsker–Varadhan-type representations. |
The key distinction is upper bound versus calibration. A surrogate can sit above the target loss everywhere and still have minimizers with bad target behavior, especially when you optimize over a restricted model class.
For binary labels \(y \in \{-1, +1\}\), any score function \(f\), the classifier \(\operatorname{sign} f\) and the logistic loss \(\phi(m) = \log(1 + e^{-m})\) with the natural logarithm,
Condition on \(X = x\). Both risks are averages over \(x\) of quantities that depend only on \(\eta(x) = P(Y = +1 \mid X = x)\) and the score \(\alpha = f(x)\). That reduces a statement about functions to a one-dimensional calculus problem, which we can solve exactly.
Step 1: the conditional risk. At a point with probability \(\eta\) and score \(\alpha\),
where \(\sigma\) is the logistic sigmoid. For \(\eta \in (0, 1)\), \(C_\eta\) is strictly convex, so its minimizer is \(\alpha^\star = \log\frac{\eta}{1 - \eta}\), which has the sign of \(2\eta - 1\). The minimum value is the binary entropy \(H(\eta) = -\eta\log\eta - (1 - \eta)\log(1 - \eta)\). At \(\eta = 0\) or \(1\) there is no minimizer: the infimum \(H = 0\) is approached as \(\alpha \to \mp\infty\), and the rest of the proof goes through with infima in place of minima.
This is calibration in one line: the best score always agrees with the Bayes decision. The rest of the proof is about how much surrogate risk a wrong decision costs.
Step 2: the price of a wrong sign. Because \(C_\eta\) is convex with its minimum on the correct side, its smallest value over wrong-sign scores is at \(\alpha = 0\), where \(C_\eta(0) = \log 2\). So wherever \(\operatorname{sign} f(x)\) is wrong,
by Pinsker's inequality, since the total-variation distance between the two coins is \(\lvert \eta - \tfrac12 \rvert\).
Pinsker appears here in its correct direction: KL controls total variation, which is what we need.
Step 3: integrate. The 0-1 excess risk is \(\mathbb{E}\big[\lvert 2\eta(X) - 1 \rvert\, \mathbf{1}\{\text{wrong sign}\}\big]\). With \(\psi(\theta) = \theta^2/2\), which is convex, Jensen's inequality and Step 2 give
The second inequality also uses \(C_\eta - H \ge 0\) at points where the sign is right. Solving \(\theta^2/2 \le R_\phi(f) - R_\phi^\star\) for \(\theta\) gives the claim. \(\blacksquare\)
Note that \(\log_2(1 + e^{-m}) \ge \mathbf{1}\{m \le 0\}\), so \(R_{01}(f) \le R_\phi(f)/\log 2\). That is true, but it bounds the absolute risk, and \(R_\phi^\star = \mathbb{E}\, H(\eta(X))\) is strictly positive whenever labels are noisy. So driving the surrogate to its minimum does not drive this bound to the Bayes risk. You need excess risk on both sides, which is what Steps 2 and 3 provide.
Repeat the proof above for the hinge loss \(\phi(m) = \max(0, 1 - m)\): compute the conditional risk \(C_\eta(\alpha)\), its minimum \(H(\eta)\), and the smallest risk with the wrong sign, \(H^-(\eta)\). What bound on the 0-1 excess risk do you get, and how does it compare with the logistic loss?
On \([-1, 1]\) the conditional risk is linear in \(\alpha\). Outside that interval one of the two terms vanishes and the other only grows.
Conditional risk. \(C_\eta(\alpha) = \eta\max(0, 1 - \alpha) + (1 - \eta)\max(0, 1 + \alpha)\). On \([-1, 1]\) this is \(1 + \alpha(1 - 2\eta)\), and it is larger outside. So for \(\eta > 1/2\) the minimum is at \(\alpha = 1\), and in general
Wrong sign. For \(\eta > 1/2\), \(C_\eta\) decreases as \(\alpha\) rises toward 0, so the best wrong-sign score is \(\alpha = 0\), with \(H^-(\eta) = C_\eta(0) = 1\). The price of a wrong sign is
Bound. That is \(\psi(\theta) = \lvert \theta \rvert\), and the same argument gives \(R_{01}(f) - R_{01}^\star \le R_\phi(f) - R_\phi^\star\): linear, where the logistic loss gave a square root. The hinge loss is calibrated with the best possible transform, but its minimizer \(\alpha = \pm 1\) carries no information about \(\eta\), so unlike the logistic loss it cannot be used to estimate probabilities.
For a binary margin \(m = y f(x)\).
For a classification-calibrated surrogate \(\phi\), Bartlett, Jordan and McAuliffe construct a non-decreasing \(\psi\) for which this holds.
For \(\eta(x) = P(Y = \cdot \mid X = x)\) and a predicted categorical distribution \(q(x)\).
The gap equals \(D_{\mathrm{KL}}\big(q(z \mid x) \,\Vert\, p(z \mid x)\big)\).
Under the usual absolute-continuity and integrability conditions.
Bartlett, Jordan & McAuliffe, Convexity, Classification, and Risk Bounds, JASA 101(473), 2006. DOI
Characterizes classification calibration and proves the ψ-transform inequality above, showing how convex surrogate regret transfers to 0-1 regret. The canonical short reference for binary surrogates.
Long & Servedio, Consistency versus Realizable H-Consistency for Multiclass Classification, ICML 2013. PMLR
Shows that Bayes consistency and consistency within a restricted class can disagree: a loss can behave well for linear scoring functions while a consistent loss fails there. The proof is by construction, and it is why "this loss is consistent" may not be the theorem you need.
Saunshi, Plevrakis, Arora, Khodak & Khandeparkar, A Theoretical Analysis of Contrastive Unsupervised Representation Learning, ICML 2019. PMLR
Under a latent-class model of positive and negative pairs, proves that low contrastive risk yields a representation with controlled average downstream classification risk. The modern pattern: build the downstream classifier explicitly and bound its loss by the self-supervised objective.
Found an error, or have a better example or a question this guide should cover? Email m.molahasani.m@gmail.com. Corrections and contributions are welcome. Last updated October 5, 2026.