Source-linked AI summary

When Does the Best Sampling Temperature Rise with the Budget? Sufficient Conditions for Pass@k

Changsu Jeong

arXiv:2608.14665v1cs.LG

TL;DR

The paper asks when the empirically observed optimal temperature increase with the pass@k budget must occur at the population level. It develops a conditional response-order theory and proves that, under this condition and strict single-peakedness, optimal temperatures are nondecreasing with budget, while a two-stratum model shows the pattern can also reverse.

  • Problem

    The paper asks which task-level response condition makes aggregate optimal temperature move weakly upward as the pass@k budget grows.

  • Method

    It orders tasks by one-sample success and analyzes normalized aggregate derivatives using monotone-likelihood-ratio power tilting toward lower-success tasks.

  • Results

    Under a nonincreasing conditional log-success response, derivative signs are nested across budgets and strictly single-peaked aggregate curves have nondecreasing optimal temperatures.

  • Takeaways & Limitations

    The theory turns the small-budget low-temperature and large-budget high-temperature pattern into a testable conditional statement that can also reverse.

  • Takeaways & Limitations

    The conclusions assume conditionally independent samples, fixed-temperature homogeneous schedules, exact verification, and a representative benchmark population, while response ordering and single-peakedness require testing.

Abstract

from arXiv · show

The temperature that maximizes pass@$k$ is often low for a small sampling budget and higher for a large budget. This pattern has been reported from Codex through recent multi-sample inference studies. It is not an algebraic property of pass@$k$: as Slocum et al. (ICLR 2025) observe, for one fixed task the maximizing temperature is independent of $k$. Building on that fixed-task observation and the hard/easy-task explanation, we give a formal population-level sufficient condition for the aggregate pattern. For task $X$, let $p_t(X)$ be one-sample success probability at temperature $t$, and define the conditional log-success response $m_t(u)=\mathbb{E}[\dot p_t(X)\mid p_t(X)=u]/u$. If $m_t(u)$ is nonincreasing in current success probability, then the normalized temperature derivative of aggregate pass@$k$ is nondecreasing in $k$. Consequently, derivative signs are nested across budgets; if each temperature-performance curve is strictly single-peaked, its unique maximizer is nondecreasing in $k$. The proof identifies the mechanism as a monotone-likelihood-ratio power tilt toward lower-success tasks. We derive a closed-form two-stratum phase diagram, including upward and downward regimes, and show that the marginal temperature derivative admits an exact $\mathrm{Beta}(2,k)$ kernel representation whose kernel concentrates at one-sample success of order $1/k$. Interpreting that scale as task-level localization additionally requires a regular, nonvanishing density-response factor near zero. A signed-moment representation yields diagnostic shape restrictions, while a short appendix records exact discrete refinements of the existing multi-configuration allocation formulation. No language model is trained, and no model query is used as an experimental measurement: the contribution is a conditional theory of an established empirical phenomenon, with assumptions that can be tested in future work.

1 Introduction

The introduction explains why aggregate pass@k optima can rise with sampling budget despite fixed-task optima being independent of k. It presents a sufficient task-response condition yielding nested derivative signs and nondecreasing optimal temperatures, alongside exact model and kernel analyses.

  • Motivation: Empirically, the best temperature often rises with sampling budget, including Codex optima near 0.2 for pass@1 and 0.8 for pass@100.The introduction also notes a rising upper hull across budgets and later qualitative replications.
  • Motivation: For a fixed task, pass@k is a strictly increasing transformation of one-sample success probability, so its maximizing temperature cannot depend on k.Benchmark-level movement instead requires heterogeneous task responses or departures from the stated sampling and decoding assumptions.
  • Main theorem: The sufficient-condition theorem gives nested temperature-derivative signs across budgets and, under strict single-peakedness, nondecreasing optimal temperatures.This is the paper’s central population-level result for the aggregate pass@k pattern.
  • Main condition: If higher temperature provides a weakly larger proportional benefit to lower-success tasks, normalized aggregate derivatives shift monotonically toward higher temperatures as k grows.The ordering arises through a monotone-likelihood-ratio weighting law, not through raw derivative magnitudes.
  • Additional analyses: The introduction previews an exact two-stratum affine phase diagram with upward and downward regimes, a Beta(2, k) kernel representation, and conditional localization analysis.It also identifies likelihood-ratio order, covariance inequalities, monotone comparative statics, and a compact moment problem as the main tools.

2 Population pass@k

Population pass@k averages the probability of at least one accepted completion across benchmark tasks, under conditional i.i.d. sampling and a perfect binary verifier. For a fixed task, pass@k rankings are budget-invariant, so aggregate budget dependence requires temperature rankings to cross across tasks.

  • Population definition: Population pass@k averages each task’s probability of at least one accepted completion under conditional i.i.d. sampling and a perfect binary verifier.For task X, one-sample success is p_t(X), with q_t = 1 − p_t(X).
  • Population definition: The population target is estimated by the usual unbiased finite-sample pass@k estimator.
  • Regularity conditions: Differential results require an interior temperature, almost-sure differentiability, integrable temperature derivatives, interchangeable differentiation and expectation, and 0 < p_t(X) < 1 almost surely.Boundary cases require separate one-sided derivatives and dominated-limit conditions.
  • Fixed-task invariance: For every fixed task x, pass@k’s temperature comparison is independent of k because 1 − (1 − p_t(x))^k is strictly increasing in p_t(x).If p_s(X) ≥ p_t(X) almost surely, then A_k(s) ≥ A_k(t) for every k.
  • Fixed-task invariance: Changing aggregate optima therefore require temperature rankings to cross across tasks, rather than arising solely because larger k rewards repeated draws.For a fixed task, output diversity affects pass@k only through the total probability mass of accepted outputs, p_t(x).

3 A monotone-temperature theorem

Under a decreasing conditional log-success response, larger sampling budgets tilt task weights toward lower one-sample success and produce nested normalized temperature-derivative signs. If aggregate pass@k curves are strictly single-peaked, their unique optimal temperatures are nondecreasing in k; these are falsifiable sufficient conditions, not universal laws.

  • Assumption 3.1: Assumption 3.1 requires the conditional log-success response m_t(u) to be nonincreasing in current one-sample success u.This is a hard-task-benefit condition: temperature has a larger proportional response, or smaller proportional cost, on lower-success tasks.
  • Budget-induced tilt: Increasing the budget induces a monotone-likelihood-ratio tilt toward tasks with lower one-sample success.The density-ratio proof leaves the factor (1 − P)^(ℓ−k), which decreases with P and shifts the law toward lower-success tasks.
  • Nested derivative signs: Theorem 3.3 establishes nested signs for the normalized temperature derivatives across budgets under the decreasing-response assumption.The result follows because m_t(P) and the budget tilt are both nonincreasing, giving them nonnegative covariance.
  • Scope and limitations: The theorem orders normalized derivatives, not raw derivative magnitudes, and the assumptions are not automatic.Without the decreasing-response condition or single-peakedness, optimal temperature may decrease, oscillate, or be nonunique.
  • Budget-monotone optimum: Strict single-peakedness of every aggregate pass@k curve implies that the unique optimal temperature t_k is nondecreasing with k.The corollary combines the nested derivative signs with unique sign-changing maxima.

4 A Beta-kernel representation of the marginal decision

The marginal temperature derivative has an exact Beta(2, k) kernel representation. The kernel concentrates near one-sample success probabilities of order 1/k, but interpreting this as task-level localization requires regular, nonvanishing density–response behavior near zero.

  • Kernel representation: The kernel has mode 1/k, mean 2/(k + 2), and concentrates on one-sample success probabilities of order 1/k.These properties describe the kernel’s sampling scale as the budget k changes.
  • Localization condition: Kernel concentration alone does not establish task-level localization at the same scale; that interpretation requires h_t to be continuous with a finite nonzero limit near zero.If h_t vanishes or varies rapidly near zero, the nonzero task-level contribution need not be centered at the kernel’s scale.

5 An exact two-stratum phase diagram

The exact two-stratum model yields a closed-form phase diagram: the optimizer is nondecreasing in k when C > 1, nonincreasing when C < 1, and constant when C = 1. An explicit C = 1/4 example demonstrates a downward sequence, ruling out any unconditional upward-temperature claim.

  • Closed-form optimizer and phase boundary: Proposition 5.1 gives a unique optimizer for k > 1 and classifies its direction: nondecreasing when C > 1, nonincreasing when C < 1, and constant when C = 1.The optimizer is obtained from the closed-form critical point and clipped to the temperature interval I.
  • Phase interpretation: C is a one-shot phase boundary: for C > 1 the optimum starts cold and moves toward the difficulty-crossing temperature as k increases, while C < 1 reverses the direction.The mechanism is that larger budgets give failures on the hard stratum greater marginal weight.
  • Worked examples: The explicit downward sequence refutes any unconditional claim that the optimal temperature must rise with budget.The easy/hard ordering reverses at t = 5/12, so Assumption 3.1 is not maintained over the relevant region.

6 A completion-level interpretation

This section interprets temperature effects through accepted-completion scores: the derivative of log success equals the accepted-set average of the completion score. Higher temperature helps when accepted completions have below-average scores and hurts when they already occupy top-score modes.

  • Completion-level score identity: Under ℓ1 differentiability and integrability, differentiating the accepted-completion sum yields an accepted-set score identity for ∂t log p_t(x).The identity is ∂t log p_t(x) = Eπ_t[∂t log π_t(Y | x) | Y ∈ C_x].
  • Gibbs-family interpretation: Higher temperature helps when accepted completions occupy lower-score modes than the current global average, and hurts when they occupy the top modes.This is the sign interpretation for the idealized global Gibbs family under the stated finiteness and differentiability assumptions.
  • Autoregressive interpretation: For autoregressive tokenwise temperature, the completion log-probability derivative sums chosen-token logit minus prefix-average logit terms with a −1/t^2 sign.Averaging this pathwise score over accepted paths recovers the completion-level identity.
  • Autoregressive interpretation: The task-level assumption is interpreted as accepted paths for currently low-success tasks being deeper in the model’s score structure.The supplied passage states this interpretation only roughly, following the tokenwise score decomposition.

7 Diagnostics for future tests

The section presents a signed-moment diagnostic: the infinite derivative sequence uniquely identifies a finite signed response measure, while nonnegative responses imply shape restrictions. Because identification requires an infinite noiseless sequence and finite-noise inversion is ill-posed, the practical output is a set of task-level diagnostics.

  • Shape test: If the temperature derivative is nonnegative almost surely, the identified moments satisfy forward-difference shape restrictions.Repeated forward differences multiply the integrand by (q−1)^r = (−p)^r.
  • Shape test: Complete monotonicity need not hold for a general signed response measure.The shape restriction therefore depends on the nonnegative-response condition rather than applying universally.
  • Limitations: Identification uses the entire infinite, noiseless derivative sequence, making inversion from finitely many noisy derivatives ill-posed.The section consequently frames the result as a diagnostic rather than a direct finite-sample identification procedure.
  • Diagnostics for future tests: The proposed practical diagnostics estimate taskwise success probabilities and nearby-temperature local log-success responses before testing their conditional response relationship.This translates the formal moment result into an empirical workflow for future tests.

8 Related work and claim boundaries

Prior work documents temperature/sample co-scaling, fixed-task invariance, pass@k reweighting, and broader adaptive or mixed-configuration inference strategies. This paper claims a narrower sufficient-condition result linking conditional log-success response to derivative nesting and ordered temperature maximizers.

  • Temperature and multi-sample inference: Chen, Du, and Chow document increasing optimal temperatures, lower temperatures at small samples, higher temperatures at larger samples, and easy/hard-task interpretations.Du et al. also report roughly single-peaked curves and propose an entropy-based selector.
  • Pass@k reweighting and inference scaling: Related work studies pass@k reweighting, direct pass@k optimization, harder-problem prioritization at larger budgets, and one-attempt success distributions.This paper specializes that distributional view to a scalar inference-time control and asks when maximizers are ordered across budgets.
  • Configuration portfolios: Configuration-portfolio work covers mixed allocation, convex relaxed failure objectives, complementary temperature-specific subsets, and prompt- and budget-conditioned decoding policies.Those settings are broader than the fixed-temperature schedules studied here.
  • Priority statement: Slocum et al. own the fixed-task invariance observation and a qualitative hard/easy-task account of aggregate movement.The paper positions its contribution as extending this observation to a population-level sufficient condition.
  • Priority statement: The paper claims only a narrow sufficient-condition result combining decreasing conditional log-success response, derivative-sign nesting, and ordered inference-time temperature maximizers.The authors report that a targeted primary-source search through July 19, 2026 found no prior result combining these elements.

9 Limitations and conclusion

The paper formalizes a sufficient condition under which aggregate pass@k favors higher temperatures as k increases, while making clear that this conclusion depends on restrictive modeling assumptions. Its mechanism is a power tilt toward tasks that still tend to fail, combined with larger conditional log-success responses for lower-success tasks.

  • Modeling commitments: The analysis assumes conditionally i.i.d. samples, fixed temperature within homogeneous schedules, exact verification, and a benchmark population matching the target.Correlation, adaptive decoding, verifier error, and distribution shift can change the objective.
  • Modeling commitments: Response order and single-peakedness must be checked rather than assumed, and finite benchmarks may introduce grid and estimation effects.These limitations bound how directly the population-level result transfers to empirical evaluations.
  • Conclusion: Pass@k leaves each fixed task’s preferred temperature unchanged but power-tilts the benchmark marginal toward tasks that still tend to fail.When lower-success tasks have larger conditional log-success responses, derivative signs become nested and the best temperature can only move upward under single-peaked curves.

Tool-use disclosure · A Exact refinements of configuration allocation · A.1 An exact integer criterion for two configurations

The appendix discloses Codex-assisted research and formalizes exact discrete consequences for two-configuration allocation, including a convexity-based criterion for strict interior mixing.

  • Tool-use disclosure: Tool-use disclosure: OpenAI Codex assisted literature search, algebraic and numerical checks, figure generation, and manuscript drafting, but is not an author.The named human author remains responsible for independently verifying the claims, references, and submitted text.
  • A Exact refinements of configuration allocation: A Exact refinements of configuration allocation: The appendix makes two discrete and asymptotic consequences of OSCA’s mixed-configuration failure objective explicit without reclaiming the general formulation.
  • A.1 An exact integer criterion for two configurations: A.1 An exact integer criterion for two configurations: For two configurations, per-sample failure probabilities qA(X) and qB(X) define success probabilities pA and pB.Conditional on X, draws are assumed independent across and within configurations.
  • A.1 An exact integer criterion for two configurations: A.1 An exact integer criterion for two configurations: With n samples from B and k −n from A, the mixed-configuration failure objective is Fk(n) = E[qA(X)k−nqB(X)n].The integer allocation ranges over n = 0, . . . , k.
  • A.1 An exact integer criterion for two configurations: A.1 An exact integer criterion for two configurations: Proposition A.1 establishes discrete convexity and a strict-mixing criterion for the integer allocation objective.
  • A.1 An exact integer criterion for two configurations: A.1 An exact integer criterion for two configurations: For k ≥2, the appendix characterizes when an interior integer schedule strictly beats both homogeneous endpoints.The proposition presents this as an if-and-only-if condition.
  • A.1 An exact integer criterion for two configurations: A.1 An exact integer criterion for two configurations: The proof derives the second difference of Fk, showing that first differences form a nondecreasing sequence.
  • A.1 An exact integer criterion for two configurations: A.1 An exact integer criterion for two configurations: Convexity implies that both strict endpoint conditions hold exactly when the minimum lies strictly below both endpoints at an interior index.

A.2 Finite task types and the large-budget limit … C Reproducibility

The appendix establishes sparse-support and maximin limits for finite task types, supplies proof details and boundary-case extensions, and documents reproducibility checks based on analytical formulas rather than experimental measurements.

  • A.2 Finite task types and the large-budget limit: The finite-task formulation assumes positive task masses, failure probabilities qji ∈(0, 1], and relaxed allocation proportions z in the simplex.Hazard coordinates are defined by aji = −log qji.
  • A.2 Finite task types and the large-budget limit: At most r configurations suffice to support an optimizer, and dominated hazard vectors can be removed without worsening the objective.The sparse-support result follows from an exposed-face argument and Carathéodory’s theorem; dominance follows because increasing hazard coordinates weakly improves the decreasing objective.
  • A.2 Finite task types and the large-budget limit: As k →∞, every accumulation point of relaxed optimizers maximizes the maximin objective m(z).This uses the uniform convergence ϕk → m on the compact simplex.
  • B.1 Association inequality used in Theorem 3.3: The association proof applies an inequality for independent identically distributed variables when both functions are nonincreasing, using f = mt and g(u) = (1 −u)ℓ−k.This is the stated application in the proof of Theorem 3.3.
  • B.2 Boundary probability cases: At boundary probabilities, tasks with pt = 0 contribute no derivative locally, while tasks with pt = 1 receive zero weight for k > 1.Extensions restrict attention to positive-weight tasks and require one-sided derivatives plus dominated convergence when taking limits.
  • C Reproducibility: The reproducibility checks verify analytical optimizers, upward and downward regimes, discrete-convexity identities, finite-difference diagnostics, and displayed numerical values.The figure is generated from analytical formulas, and the computations serve as regression checks rather than substitutes for the proofs.
Loading 2608.14665v1…