Question 3

Is objective A a valid surrogate for B?

Your algorithm optimizes a convenient loss or divergence A. Does a small A provably mean a small error in the metric B you actually care about?

Symptoms

Is this your question?

Sub-cases

Which version are you facing?

Case Usually start with
Is a pointwise upper bound enough? Direct domination, for example the 0-1 loss is at most the hinge loss.
Need Bayes consistency, not just an upper bound? Conditional-risk analysis and the calibration (ψ-transform) inequality.
A divergence stands in for a distributional discrepancy? Pinsker, data processing, variational representations, proper-scoring identities.
A self-supervised objective should imply downstream performance? Latent-variable or augmentation assumptions, an explicit downstream classifier, and a transfer bound.
A variational lower or upper objective? Jensen, Fenchel duality, Donsker–Varadhan-type representations.
The recipe

Condition on the input, then integrate

  1. Condition on \(X = x\) wherever you can. Write the conditional target risk and conditional surrogate risk as functions of \(\eta(x) = P(Y \mid X = x)\).
  2. Characterize each conditional minimizer.
  3. Prove pointwise domination, or a conditional regret-transfer inequality.
  4. Integrate over \(X\).
  5. Only then add estimation and optimization error. Schematically, with the exact form set by your calibration result,
\[R_B(\hat f) - R_B^\star \;\le\; \Psi^{-1}\Big( \underbrace{R_A(\hat f) - \inf_{f \in \mathcal{F}} R_A(f)}_{\text{estimation + optimization}} + \underbrace{\inf_{\mathcal{F}} R_A - R_A^\star}_{\text{approximation}} \Big).\]

The key distinction is upper bound versus calibration. A surrogate can sit above the target loss everywhere and still have minimizers with bad target behavior, especially when you optimize over a restricted model class.

0.70
The Bayes decision is +1 when \(\eta > 1/2\).
The lossesevery one here sits above the 0-1 loss
Conditional risk at this \(\eta\)\(C_\eta(\alpha) = \eta\,\phi(\alpha) + (1 - \eta)\,\phi(-\alpha)\)
0-1 loss; thick bar marks the best score selected surrogate best risk with the wrong sign
Calibration, one input at a time. The recipe says to condition on \(X = x\). Here that leaves a single number \(\eta\), and the surrogate is calibrated when its best score always has the sign of \(2\eta - 1\) with a strict penalty for the wrong sign. The last loss upper-bounds the 0-1 loss everywhere, yet its conditional risk is minimized at a score of 0 for every \(\eta\): an upper bound is not a guarantee. The logistic loss is drawn in base-2 units, \(\log_2(1 + e^{-m})\), so that it sits above the 0-1 loss; the proof below uses the natural logarithm, which only rescales it. Scores are searched on \([-3, 3]\), so near \(\eta = 0\) or \(1\) a best score at the edge stands for an optimum that is really at infinity.
A complete proof

The logistic loss is calibrated, with a rate

The claim

For binary labels \(y \in \{-1, +1\}\), any score function \(f\), the classifier \(\operatorname{sign} f\) and the logistic loss \(\phi(m) = \log(1 + e^{-m})\) with the natural logarithm,

\[R_{01}(f) - R_{01}^\star \le \sqrt{2\,\big(R_\phi(f) - R_\phi^\star\big)}.\]
Assumptions
  • \(f\) ranges over all measurable functions, so \(R_\phi^\star\) is the unrestricted minimum.
  • At a score of exactly 0 the classifier predicts the label that disagrees with the Bayes decision. This adversarial tie-breaking only makes the claim harder.
First move, and why

Condition on \(X = x\). Both risks are averages over \(x\) of quantities that depend only on \(\eta(x) = P(Y = +1 \mid X = x)\) and the score \(\alpha = f(x)\). That reduces a statement about functions to a one-dimensional calculus problem, which we can solve exactly.

Show the proof, with commentary

Step 1: the conditional risk. At a point with probability \(\eta\) and score \(\alpha\),

\[C_\eta(\alpha) = \eta \log(1 + e^{-\alpha}) + (1 - \eta)\log(1 + e^{\alpha}), \qquad C_\eta'(\alpha) = \sigma(\alpha) - \eta,\]

where \(\sigma\) is the logistic sigmoid. For \(\eta \in (0, 1)\), \(C_\eta\) is strictly convex, so its minimizer is \(\alpha^\star = \log\frac{\eta}{1 - \eta}\), which has the sign of \(2\eta - 1\). The minimum value is the binary entropy \(H(\eta) = -\eta\log\eta - (1 - \eta)\log(1 - \eta)\). At \(\eta = 0\) or \(1\) there is no minimizer: the infimum \(H = 0\) is approached as \(\alpha \to \mp\infty\), and the rest of the proof goes through with infima in place of minima.

This is calibration in one line: the best score always agrees with the Bayes decision. The rest of the proof is about how much surrogate risk a wrong decision costs.

Step 2: the price of a wrong sign. Because \(C_\eta\) is convex with its minimum on the correct side, its smallest value over wrong-sign scores is at \(\alpha = 0\), where \(C_\eta(0) = \log 2\). So wherever \(\operatorname{sign} f(x)\) is wrong,

\[C_\eta\big(f(x)\big) - H(\eta) \ge \log 2 - H(\eta) = D_{\mathrm{KL}}\big(\mathrm{Bern}(\eta) \,\Vert\, \mathrm{Bern}(\tfrac12)\big) \ge \frac{(2\eta - 1)^2}{2},\]

by Pinsker's inequality, since the total-variation distance between the two coins is \(\lvert \eta - \tfrac12 \rvert\).

Pinsker appears here in its correct direction: KL controls total variation, which is what we need.

Step 3: integrate. The 0-1 excess risk is \(\mathbb{E}\big[\lvert 2\eta(X) - 1 \rvert\, \mathbf{1}\{\text{wrong sign}\}\big]\). With \(\psi(\theta) = \theta^2/2\), which is convex, Jensen's inequality and Step 2 give

\[\psi\big(R_{01}(f) - R_{01}^\star\big) \le \mathbb{E}\Big[\psi\big(\lvert 2\eta - 1 \rvert\big)\, \mathbf{1}\{\text{wrong}\}\Big] \le \mathbb{E}\big[C_\eta(f(X)) - H(\eta(X))\big] = R_\phi(f) - R_\phi^\star .\]

The second inequality also uses \(C_\eta - H \ge 0\) at points where the sign is right. Solving \(\theta^2/2 \le R_\phi(f) - R_\phi^\star\) for \(\theta\) gives the claim. \(\blacksquare\)

A tempting approach that fails

Note that \(\log_2(1 + e^{-m}) \ge \mathbf{1}\{m \le 0\}\), so \(R_{01}(f) \le R_\phi(f)/\log 2\). That is true, but it bounds the absolute risk, and \(R_\phi^\star = \mathbb{E}\, H(\eta(X))\) is strictly positive whenever labels are noisy. So driving the surrogate to its minimum does not drive this bound to the Bayes risk. You need excess risk on both sides, which is what Steps 2 and 3 provide.

If you relax an assumption
  • A restricted class \(\mathcal{F}\): replace \(R_\phi(f) - R_\phi^\star\) by estimation error plus the approximation term \(\inf_{\mathcal{F}} R_\phi - R_\phi^\star\).
  • Multiclass: softmax cross-entropy is calibrated, but consistency within a restricted class is subtler; see Long & Servedio below.
  • A low-noise condition on \(\eta\) improves the exponent from \(1/2\) toward 1 (Bartlett, Jordan & McAuliffe).
Try it yourself

The hinge loss, by the same recipe

The exercise

Repeat the proof above for the hinge loss \(\phi(m) = \max(0, 1 - m)\): compute the conditional risk \(C_\eta(\alpha)\), its minimum \(H(\eta)\), and the smallest risk with the wrong sign, \(H^-(\eta)\). What bound on the 0-1 excess risk do you get, and how does it compare with the logistic loss?

Hint

On \([-1, 1]\) the conditional risk is linear in \(\alpha\). Outside that interval one of the two terms vanishes and the other only grows.

Show a solution

Conditional risk. \(C_\eta(\alpha) = \eta\max(0, 1 - \alpha) + (1 - \eta)\max(0, 1 + \alpha)\). On \([-1, 1]\) this is \(1 + \alpha(1 - 2\eta)\), and it is larger outside. So for \(\eta > 1/2\) the minimum is at \(\alpha = 1\), and in general

\[H(\eta) = 1 - \lvert 2\eta - 1 \rvert.\]

Wrong sign. For \(\eta > 1/2\), \(C_\eta\) decreases as \(\alpha\) rises toward 0, so the best wrong-sign score is \(\alpha = 0\), with \(H^-(\eta) = C_\eta(0) = 1\). The price of a wrong sign is

\[H^-(\eta) - H(\eta) = \lvert 2\eta - 1 \rvert.\]

Bound. That is \(\psi(\theta) = \lvert \theta \rvert\), and the same argument gives \(R_{01}(f) - R_{01}^\star \le R_\phi(f) - R_\phi^\star\): linear, where the logistic loss gave a square root. The hinge loss is calibrated with the best possible transform, but its minimizer \(\alpha = \pm 1\) carries no information about \(\eta\), so unlike the logistic loss it cannot be used to estimate probabilities.

The toolkit

What each tool gives you, and what it costs

Pointwise domination

\[\ell_B(z) \le C\,\ell_A(z)\ \ \forall z \;\Longrightarrow\; R_B(f) \le C\,R_A(f), \qquad \text{e.g. } \mathbf{1}\{m \le 0\} \le (1 - m)_+\]

For a binary margin \(m = y f(x)\).

Gives
An immediate risk upper bound.
Costs
A true pointwise inequality.
Wrong tool when
What matters is excess risk or calibration, not the absolute risk.

Calibration and the ψ-transform

\[\psi\big(R_{01}(f) - R_{01}^\star\big) \le R_\phi(f) - R_\phi^\star\]

For a classification-calibrated surrogate \(\phi\), Bartlett, Jordan and McAuliffe construct a non-decreasing \(\psi\) for which this holds.

Gives
Converts surrogate excess risk into target excess risk.
Costs
Population conditional-risk calibration; the exact \(\psi\) depends on \(\phi\).
Wrong tool when
You only know \(\phi \ge \ell_{01}\); that alone gives no useful regret-transfer rate.

Proper log-loss identity

\[R_{\log}(q) - R_{\log}(\eta) = \mathbb{E}_X\, D_{\mathrm{KL}}\big(\eta(X)\,\Vert\, q(X)\big)\]

For \(\eta(x) = P(Y = \cdot \mid X = x)\) and a predicted categorical distribution \(q(x)\).

Gives
An exact characterization of excess risk; uniqueness up to zero-probability labels.
Costs
Well-defined probabilities and finite cross-entropy.
Wrong tool when
Scores are unnormalized, or the metric you care about has no controlled relation to conditional KL.

Pinsker's inequality

\[\lVert P - Q \rVert_{\mathrm{TV}} \le \sqrt{\tfrac12 D_{\mathrm{KL}}(P \Vert Q)}\]
Gives
KL control implies control of every bounded test function and event probability.
Costs
\(P \ll Q\) for finite KL.
Wrong tool when
You try to bound KL by total variation; that direction needs stronger assumptions.

Jensen and the ELBO

\[\log p(x) = \log \mathbb{E}_q\!\left[\frac{p(x, z)}{q(z \mid x)}\right] \ge \mathbb{E}_q \log \frac{p(x, z)}{q(z \mid x)}\]

The gap equals \(D_{\mathrm{KL}}\big(q(z \mid x) \,\Vert\, p(z \mid x)\big)\).

Gives
A tractable lower bound on the log-likelihood.
Costs
Support and integrability conditions.
Wrong tool when
An ELBO increase is read as an improvement in an unrelated task metric.

Donsker–Varadhan variational form

\[\log \mathbb{E}_Q\, e^{f} = \sup_{P \ll Q} \big\{ \mathbb{E}_P f - D_{\mathrm{KL}}(P \Vert Q) \big\}\]

Under the usual absolute-continuity and integrability conditions.

Gives
Turns KL divergences and partition functions into variational objectives.
Costs
Exponential integrability and absolute continuity.
Wrong tool when
The bias and variance of the finite-sample estimator dominate the population identity.
Worked examples

How published papers answer it

  1. Bartlett, Jordan & McAuliffe, Convexity, Classification, and Risk Bounds, JASA 101(473), 2006. DOI

    Characterizes classification calibration and proves the ψ-transform inequality above, showing how convex surrogate regret transfers to 0-1 regret. The canonical short reference for binary surrogates.

  2. Long & Servedio, Consistency versus Realizable H-Consistency for Multiclass Classification, ICML 2013. PMLR

    Shows that Bayes consistency and consistency within a restricted class can disagree: a loss can behave well for linear scoring functions while a consistent loss fails there. The proof is by construction, and it is why "this loss is consistent" may not be the theorem you need.

  3. Saunshi, Plevrakis, Arora, Khodak & Khandeparkar, A Theoretical Analysis of Contrastive Unsupervised Representation Learning, ICML 2019. PMLR

    Under a latent-class model of positive and negative pairs, proves that low contrastive risk yields a representation with controlled average downstream classification risk. The modern pattern: build the downstream classifier explicitly and bound its loss by the self-supervised objective.

Common traps

Where proofs of this kind go wrong

Resources

Where to go next

← Back to the guide

Found an error, or have a better example or a question this guide should cover? Email [email protected]. Corrections and contributions are welcome. Last updated October 5, 2026.