X-ray sensitivity calculations commonly replace the Cash-statistic
likelihood-ratio improvement,
,
by a single expected value. Near a detection threshold, the missing
quantity is its conditional width. We derive explicit pixel-sum
expressions for the leading Asimov location and Fisher-projected
variance of
in a three-parameter PSF fit, conditioned on the fitted amplitude, local
background, pixelized PSF, and fitting mask. The mean extends earlier
analytic sensitivity calculations; the leading projected variance is the
central new theoretical result. At low signal we write the
observable-conditioned distribution as the location–scale kernel
,
where
identifies the fitter and instrumental configuration. Across
810,987 clean XMM-Newton PN band-4 reference-fitter
realizations, the 266,837 high-signal fits have a
standardized-residual mean of
and width of
.
Compact, separately fitted functions of
describe the low-signal shift and width recovery. In out-of-sample
threshold tests at
,
8, and 10, adding both corrections reduces the mean passing-fraction
error from 0.0136 for raw Fisher theory to 0.0020, close to the 0.0020
empirical-binned benchmark. Whole-position validation rejects deriving
the width correction from the mean correction alone. A targeted paired
simulation identifies centroid relocation and position maximization as
the largest positive contribution to the low-signal shift in the
integer-grid reference fitter. A standalone emldetect case
study shows how the analytic backbone can be supplemented by a sparse
position-dependent mean calibration. Standalone M1/M2
emldetect development simulations also show that the PN
reference-fitter coefficients do not transfer directly: the
location–scale architecture is reusable, but its coefficients are
configuration specific. The kernel supports fitted-amplitude
threshold-crossing maps and catalog-space selection probabilities.
True-flux completeness and Eddington correction additionally require the
joint model
,
a source-population prior, and normalization by the full selection
event.
Source detection in Poisson-count photon imaging is commonly carried
out by fitting a parametric source model on top of a known background
and summarizing the result by a likelihood-ratio statistic. Under the
standard Cash-statistic formulation (Cash
1979), the natural statistical object is the Cash improvement
,
with
proportional to twice the negative Poisson log-likelihood. Detection
pipelines typically report a monotone transform of
as a “detection likelihood”; in the XMM-Newton Science Analysis System,
for example, emldetect (Cruddace et al. 1988; Watson et al. 2009) fits a
three-parameter source model (one free amplitude and two free position
coordinates) to Poisson-count images and reports
,
an incomplete-gamma transform of the underlying
.
The reported detection-likelihood statistic is operationally important:
it is used to decide which sources enter a catalog, to define detection
thresholds, and to construct sensitivity maps.
Sensitivity maps and source counting methodologies have been
extensively developed and applied in major X-ray surveys (Cappelluti et al.
2009; Georgakakis et al. 2008; Laird et al. 2009; Luo et al. 2017, e.g.,; Mateos et al. 2008; Puccetti et al. 2009;
Wang et al. 2016; Xue et al. 2011).
The statistical treatment of Poisson source detection, limits, and
likelihood ratio methods also has a long history in high-energy
astrophysics (Broos et al. 2010, e.g.,;
Gehrels 1986; Kashyap et al. 2010; Kraft et al. 1991; Protassov et al.
2002; Starck et al.
2002). The emldetect tool has been the standard
for the XMM-Newton serendipitous survey catalogs (Carrera et al. 2007; Rosen et al. 2016; Traulsen et al. 2019;
Watson et al. 2009;
Webb et al.
2020).
We use strictly as the operational monotone threshold scale defined by SAS, not as a calibrated frequentist false-alarm probability. The usual null reference assumes interior regularity, although source amplitude lies on a boundary and position is unidentifiable under the null (Chernoff 1954; Protassov et al. 2002; Self & Liang 1987). Monte Carlo experiments by Stewart (2009) found that the corresponding Cash-statistic reference can nevertheless be a good practical approximation under representative X-ray source-search conditions after negative-amplitude solutions are excluded. The low-count Cash statistic can also depart from its high-count limit (Bonamente 2020). We do not attempt to resolve that null-calibration debate. The conversion between and is used only to translate a chosen pipeline threshold into the statistic whose conditional distribution is modeled below.
The practical threshold regime is low signal. For many XMM-Newton applications the relevant sensitivity-map range is of order 6–10. This is the regime where source counts are steep, Poisson fluctuations matter, and a small shift in the detection threshold can change the inferred limiting flux. It is therefore also the regime where Eddington bias becomes important: random upward fluctuations are preferentially selected near the detection boundary, and the selected sample no longer represents the untruncated parent distribution (Eddington 1913; Wang 2004).
The analytical mean is not the missing result. Stewart (2009) derived an amplitude-only
Cash-statistic expectation and proposed its use in sensitivity
calculations, extending earlier matched-filter sensitivity work (Stewart 2006). SAS esensmap
likewise uses an emldetect-style analytical Cash model,
following the 3XMM and 4XMM pipelines (Rosen et al. 2016; Traulsen et al. 2019); the New-ANGELS
XMM survey evaluated a related expectation at the fitted rather than
injected flux (Huang et al. 2025,
appendix). Moments of the Cash statistic have also been derived
for spectral goodness-of-fit problems (Kaastra 2017), a different estimand from
the source-versus-background improvement studied here.
The unresolved quantity in these sensitivity calculations is the conditional width of the source-detection improvement. Among the X-ray sensitivity-map methods reviewed here, we find no previous explicit pixel-sum variance for after the amplitude and two position directions have been fitted, nor a threshold calculation that propagates that width in fitted-observable space. We therefore separate the problem into an analytical location and scale and an empirical low-count correction. The Asimov/Kullback–Leibler term supplies the location, while a Fisher/Wald projection supplies the leading variance (Cowan et al. 2011; Wilks 1938). This is the pixelized Poisson-imaging counterpart of related Fisher analyses of joint photometry and astrometry (Mendez et al. 2014).
The resulting model has a common architecture but not a universal set of empirical coefficients: Here and denote the pixelized PSF and fitting mask, and denotes the effective fitter/search configuration. The functions and describe the low-count location and width corrections. They must be estimated independently: a correct mean is needed to remove within-bin gradients before measuring the width, but it does not determine that width. The nonlinear incomplete-gamma transform from to is applied only after this distribution is constructed.
Section 2 fixes the observable-conditioning and selection rules.
Section 3 derives the analytical location and Fisher-projected variance.
Section 4 validates the reference-fitter kernel, including separate
low-count corrections to the mean and width. Section 5 retains
standalone emldetect as a pipeline-specific calibration
case study and quantifies the cross-configuration transfer boundary.
Section 6 develops fitted-amplitude threshold-crossing and
catalog-inference consequences. Sections 7 and 8 state the evidence
tiers and conclusions.
The central statistical object is the conditional distribution of in fitted-observable space,
where is the fitted source amplitude, is the local background in counts per pixel, is the pixelized PSF template used by the fitting model, is its image support, and records the remaining search and optimizer convention. Section 3 derives the leading location and variance. Section 4 measures the low-count corrections in a matched reference fitter, Section 5 examines a standalone pipeline calibration, and Section 6 states the uses and additional assumptions required for a selection function.
In simulations, is a generation label and is the appropriate conditioning variable for a separate generative selection model . It is not observed for catalog sources and is not the conditioning variable of the fitted-observable kernel developed here. We therefore use only for generative diagnostics and condition all moment estimates on ; true-flux completeness is treated separately in Section 6. A result conditioned on tests behavior at a known injected flux, whereas a result conditioned on tests the kernel after a source has been fit. Only the latter is the estimand of the observable-conditioned kernel and fitted-amplitude crossing maps studied here.
For each source or fake-source trial the theory prediction is evaluated at the observed fitted quantities,
and the residual is
For a sample or bin,
Only when is approximately constant across the bin does the observed variance reduce to ; mixed- or mixed-position samples otherwise introduce apparent variance effects that are not properties of the conditional theory. We therefore estimate a local conditional mean in the same cells used for the width and compute the standard deviation only after subtracting that local trend. The mean is essential to this step, but the remaining width is an independent statistical target. Binning by instead mixes fitted amplitudes and can inflate the apparent position dependence. Binning after a cut on observed , , or the residual is worse: it truncates the distribution whose moments are being estimated. Neither operation is used in the validation below.
is a nonlinear monotone tail-probability transform of ,
with and for the single-source three-parameter fit. The corresponding inverse is where is the regularized upper incomplete gamma function and denotes inversion with respect to its second argument. Thus correspond to , respectively. A Gaussian residual in is therefore not Gaussian in , and selecting on observed is a selection on the same noisy quantity whose distribution is being measured: it truncates the residual distribution and biases its mean upward near the threshold. Threshold ranges are accordingly translated into using the three-parameter convention and applied as theoretical or predicted conditions (such as $\mu_{3p,\rm ref}\in[\Delta C_\mathrm{ML6},\Delta C_\mathrm{ML10}]$) or as bins whose typical predicted lies in that interval. This keeps the conditioning independent of the observed stochastic fluctuation in .
The Kullback–Leibler/Asimov construction for the location and the Fisher/Wald construction for local fluctuations are standard likelihood tools (Cowan et al. 2011; Wilks 1938). Their evaluation for a pixelized Poisson source fit yields the specific result needed here: an explicit variance after the amplitude and both position score directions have been removed. We derive the location and this projected variance for one free amplitude and two free coordinates. The variance, rather than the analytical mean alone, is the main theoretical contribution.
Consider a fitting region containing pixels indexed by . The observed counts are independent Poisson random variables,
For a single source on a locally known background, the fitted source model is
Here is the background expectation in pixel , is the source amplitude in the same units used by the image model, and is the pixelized PSF template shifted to position . The PSF is normalized over the full template grid, ; the mask then selects the fitting footprint. Thus denotes the total source amplitude in the template convention, not the counts enclosed by the fitting mask. In the current PN band-4 validation runs, is usually a constant inside the fitting region, but the notation allows the background to vary by pixel.
A binary fitting mask selects the pixels included in the fit; the primary results in this paper use binary masks. A weighted-mask extension exists for fractional weights, but when the weights are exactly 0 or 1 the weighted formula reduces to the binary-mask convention used here.
Ignoring constants independent of the model parameters, the Cash statistic is
The null model is the background-only model, . The likelihood-ratio Cash improvement is
With this sign convention, a better source fit gives positive . At a fixed fitted model , the statistic can be written as
The observable-conditioned theory evaluates the moments at the fitted result
The Asimov image (the expected dataset where observed counts equal the model expectation; Cowan et al. 2011) for this fitted model is . Substitution into the Cash improvement gives the leading conditional location
For a constant background and normalized PSF , this becomes the explicit pixel sum:
This mean depends on the pixelized PSF and the mask through the full pixel sum. It is not a function only of a Gaussian width or of low-order PSF moments, although such summaries can be useful diagnostics.
If position is fixed and only the amplitude is fitted, Equation ([eq:mu3p]) reduces, up to notation and the treatment of a small empirical offset, to the Cash-statistic expectation used by Stewart (2009). The extensions needed here are conditioning on the reported fitted amplitude, carrying the fitting mask explicitly, and propagating the two fitted position directions into the variance.
Writing , Eq. [eq:mu3p] becomes , which is twice the per-pixel Poisson Kullback–Leibler divergence from background to source plus background, summed over the fitting region. It is the Asimov location backbone for the local likelihood-ratio approximation (Cowan et al. 2011), not a claim that the finite-count, free-position conditional distribution is exactly noncentral . The convexity of guarantees .
An important asymmetry exists between the mean and the variance derived below. The mean is a zeroth-order Asimov quantity: it substitutes the expected counts into the Cash improvement but does not account for the absorption of Poisson fluctuations by the fitted parameters (the finite-search/argmax behavior associated with position optimization). The variance , by contrast, already includes a first-order correction for this absorption via the Fisher projection of Section 3.6. This asymmetry motivates separate empirical functions for location and scale at low fitted signal. Position parameters can migrate toward a local upward fluctuation, an argmax behavior structurally analogous to a finite look-elsewhere effect (Gross & Vitells 2010); the paired diagnostic in Section 4 measures this mechanism only for its bounded integer-grid reference fitter. Pipeline-specific effects are kept separate in Section 5.
The best-fit parameters satisfy the score equations. For any fitted parameter ,
Linearizing around the fitted model, write
The first-order score constraint is
There are three such constraints:
These constraints are the reason a three-parameter fit is not equivalent to a simple unconstrained Poisson sum. Fluctuations along the fitted amplitude and position directions are absorbed by the fit and must be projected out when computing the residual variance.
At fixed fitted model, the first-order stochastic part of is
If the fitted parameters were not constrained by the score equations, the variance of this linear residual would be
This term, , is useful as a diagnostic but is not the final variance formula for the fitted likelihood ratio. It ignores the fact that , , and were chosen to maximize the likelihood. Because the Fisher projection below subtracts a positive semidefinite quadratic form, is an upper bound on the projected variance within the local Gaussian approximation.
Define the Fisher information matrix for the three fitted parameters:
For , the entries are
The covariance between the residual and the score direction is encoded by
Explicitly,
Under the Gaussian/Fisher approximation, conditioning on the fitted score constraints subtracts the projection of the residual onto the fitted parameter subspace. This is the imaging form of the usual Fisher/Wald projection and follows from the multivariate Gaussian conditional-variance identity, where denotes the vector of linearized score constraints. The resulting variance is
Equation ([eq:var3p]) is the central analytic result. The first term is the variance of the fixed-model linear residual. The second removes the part of that fluctuation absorbed when amplitude and position are re-estimated. Both terms are evaluated from the fitted source model, pixelized PSF, background, and mask; no injected amplitude appears.
The projection can be computed conveniently via the block structure of the Fisher matrix. Writing where contains the -position cross terms and is the position block for and , the projection becomes where . This block form makes the position projection explicit and is numerically convenient for a large grid of and values. If the position directions are removed and only is fit, the formula reduces to the amplitude-only one-parameter expression.
The variance expression is a local Gaussian/Fisher approximation around the fitted model. It is expected to be accurate when the likelihood surface is close to quadratic and when the fitted source is strong enough that position optimization does not absorb isolated Poisson noise features. The properly -conditioned variance diagnostic, and the global low- correction that absorbs the residual deficit, are presented in Section 4.3 and Figure [fig:deltac-scatter-variance]. These Fisher/Wilks approximations fail at the boundary and become incomplete in the low-signal threshold band; the remaining threshold-band location and width corrections are therefore measured empirically in Section 4. Equation ([eq:var3p]) is not, by itself, a validated tail model at –10.
The reference fitter is not intended to be a reimplementation or
replacement of SAS emldetect. Instead, it provides a
controlled matched-model baseline: the simulated sources, the Poisson
likelihood fitter, and the theoretical moment calculation all use the
same psfgen PSF template, fitting footprint, and likelihood
convention. This construction separates two questions that are otherwise
entangled in the standalone pipeline comparison. First, it tests whether
the analytic observable-conditioned
mean and Fisher-projected variance close when the data-generating model,
fitting model, and theory are matched. Second, by comparing standalone
emldetect against this controlled baseline, we can identify
the remaining discrepancy as a combination of low-signal asymptotic bias
and pipeline-specific implementation residuals (optimizer behavior, PSF
sampling, numerical discretization, and mask conventions). The matched
experiment is therefore the primary test of Equation ([eq:var3p]) and of the low-count
location–scale form. Standalone pipeline transfer is evaluated
separately in Section 5.
The reference validation uses high-statistics simulations for the XMM-Newton PN band-4 configuration, spanning 17 detector positions, eight injected source strengths, six background levels, and 1000 realizations per grid cell. The design contains 816,000 fits. The location–scale analysis retains 810,987 clean rows after requiring , , finite moments, and successful fits.
The pixelized PN band-4 PSF templates were generated with
psfgen under XMM-Newton SAS v20.0.0. The matched source
fits do not call emldetect: they use a frozen integer-grid
position search with bounded scalar amplitude optimization and an
circular mask. No detector position is removed from the reference-fitter
validation. The standalone calibration grid in Section 5 excludes an
edge-affected calibration position whose local PSF geometry violates the
stationary-background assumption used by the empirical residual model.
For standalone comparisons, the empirical statistic is extracted
directly from the catalog
column and converted using
with
.
Each simulated source is fit with a three-parameter Poisson likelihood fitter. For every row, the theory is evaluated at the fitted source strength:
The diagnostic residual is
This is an observable-conditioned validation because the prediction uses the fitted amplitude , not the injected amplitude .
Figure [fig:deltac-scatter-variance] is the central raw-data diagnostic for the matched experiment. It shows three facts needed by the distributional model: the pixel-sum location follows the dominant mean trend, the remaining low-signal location shift is smooth, and the Fisher-projected scale approaches the observed width at high signal. The rows show three representative detector radii and the colors show five backgrounds. The right column is the key variance diagnostic: the standard deviation is measured from rather than from raw , so gradients of the mean inside a bin do not masquerade as additional width. The dashed curves are descriptive fits to the same PN band-4 sample. The normalized and parameterization used for validation is introduced below.

Table [tab:ref-fitter-validation] gives the two anchor regimes. At high signal (), the standardized residual has mean 0.068 and width 0.977. Thus both the Asimov location and the projected scale close in the matched model. At , the mean shifts upward to 0.922 Fisher units while the width contracts to 0.856. A low-count model must therefore correct location and scale separately.
lrrr
& 137,310 &
& 0.856
& 266,837 &
& 0.977
Define the raw normalized residual The low-count calibration estimates its location and scale as functions of the observable expected likelihood strength . The preferred compact functions are with and The conditional location and width are therefore The weighted RMSE values of and over 32 bins are 0.0105 and 0.0107, respectively. Both functions approach their Fisher limits, and , as increases. The low-count location shift qualitatively recalls the empirical additive offset discussed by Stewart (2009), but here it is neither constant nor interpreted as the same correction.

The corrections were scored by threshold-passing probabilities at the fixed values corresponding to , 8, and 10. No row was selected by observed , , or a residual. In the reconstructed realization-block split, the mean absolute error over the three thresholds falls from 0.013585 for the uncorrected Fisher kernel to 0.002428 after the location correction and to 0.001999 after both location and width corrections. The last value is close to the 0.001967 empirical-binned benchmark. Leave-one-position-out and leave-one-background-out tests give the same ordering (Table [tab:kernel-oos]).
lrrrr Realization block & 0.013585 & 0.002428
& 0.001999 & 0.001967
Leave one position out & 0.013641 & 0.002768 & 0.002276
&
Leave one background out & 0.013579 & 0.002520 & 0.002106
&
The improvement from the location correction is the largest, but the width term gives a reproducible further gain and nearly reaches the non-parametric cell benchmark. The result is a calibration of threshold-passing probabilities in the matched reference fitter. It is not evidence that transformed values are Gaussian, nor an injection-completeness measurement at fixed .
Both and vary with , creating a visually strong pooled association. We tested whether an affine function of the fitted mean correction could replace the independently fitted . Under 17 whole-position folds, the independent and mean-linked width models give RMSE 0.040904 and 0.045494. Their ratio, 1.1122, misses the frozen 1.10 compression criterion. After the common trend is removed, residual Spearman correlations between location and width are overall and for . The mean must be modeled to remove gradients before estimating a width, but the width requires its own correction function. This grouped test reuses the historical reference sample and is a model-discrimination result, not an independent confirmation data set.
A targeted paired rerun saved oracle, amplitude-only,
relocated-position, and full likelihood components for
96,000 fits across four detector positions. Among
49,193 clean rows with
,
amplitude optimization contributes
to the mean raw
budget, whereas centroid relocation and position maximization contribute
.
Negative oracle/noise and fitted-conditioning terms cancel much of this
gain, leaving a total fitted shift of
.
This supports a finite-position argmax interpretation in that
integer-grid reference fitter, structurally analogous to look-elsewhere
behavior (Gross & Vitells
2010). It does not establish the same component budget for
continuous optimization or standalone emldetect.
Within the PN band-4 matched configuration, the compact kernel leaves RMS location trends of 0.0073 across positions and 0.0377 across backgrounds. This within-configuration stability is useful but does not imply coefficient universality. PSF preparation, fitting mask, search domain, and optimizer are part of in Equation ([eq:master-kernel]); the cross-configuration test in Section 5 shows that changing them can change both the sign of and the scale .
The high-signal closure alone would not justify threshold use: for the current PN band-4 configuration, –10 lies mostly below . The out-of-sample tests in Section 4.4 address this gap empirically for the matched reference fitter. They validate the compact location–scale kernel over the tested threshold range, but not an exact Gaussian tail law at arbitrarily small signal and not the behavior of another optimizer or PSF convention.
End-to-end fake-source injection remains the appropriate route for an
absolute survey completeness calibration. It includes candidate
generation, source confusion, background-map construction, and every
pipeline cut, as in modern simulation-calibrated survey analyses (Brunner et al.
2022; Liu et al.
2022). The purpose of the analytical backbone is different.
Equations ([eq:mu3p]) and ([eq:var3p])
carry the continuous dependence on background, PSF, and mask. A matched
simulation then calibrates two dimensionless functions,
and
,
rather than tabulating a dense distribution at every position and flux.
Section 5 tests how much additional state is needed when the fitter is
changed to standalone emldetect.
The matched reference fitter isolates the statistical theory, but SAS
emldetect has a different optimizer, internally rendered
PSF, search surface, and mask convention. This section retains the
historical PN band-4 standalone experiment as a case study of the
required extra calibration. Its first result is a position-dependent
calibration of the mean crossing; it is not an absolute
completeness measurement. The second result is a transfer test showing
that neither this calibration nor the reference-fitter
and
coefficients can be assumed to hold for another camera/band/search
configuration.
The raw theory supplies the Asimov location
If standalone emldetect exactly matched the same model, the
threshold limit could be found from
In practice, this uncalibrated-theory baseline is only a borderline
approximation in the threshold band. In the validation, it has maximum
errors of 19.9% at
,
15.6% at
,
and 13.3% at
.
The error is position-dependent rather than a single scalar offset.
Figure [fig:emldetect-scatter-variance]
is the central diagnostic for the operational
standalone-emldetect chain. It is deliberately parallel to
Figure [fig:deltac-scatter-variance]:
the same observable-conditioned structure remains visible, but the
standalone pipeline introduces a position-dependent residual relative to
the reference-fitter theory. This is the visual reason that the final
calibration is not the raw analytic equation, but the analytic equation
plus an empirical
calibration.

We therefore define the standalone emldetect
residual
Here
is the standalone emldetect Cash improvement converted with
the single-image
convention, and
is the fitted emldetect source amplitude. The calibrated
threshold equation is
The sign follows directly from the residual definition. If
,
standalone emldetect produces larger
than the raw theory at the same fitted source strength, so the
theoretical threshold needed for the sensitivity limit is lower. If
,
the required source amplitude is higher.
The calibration uses paired fake-source simulations in which each
realization is analyzed both by the reference fitter and by standalone
SAS emldetect. The current calibration grid has 14 PN
band-4 calibration positions after excluding the edge-affected case. The
target threshold band is defined in
space using the threshold conversions of Equation ([eq:detml-inverse]). Rows are
selected by the reference-theory prediction:
This is important: the threshold sample is not selected by observed
or by observed
,
because those quantities carry the stochastic fluctuation and
implementation residual we are trying to calibrate. The selection is
based on the predicted/reference
in the observable-conditioned theory. The baseline calibration grid is
built at
;
background transfer to
and
is included in the uncertainty budget below.
The residual is estimated at each detector position from the threshold-band simulation sample. We compare a 2D interpolation over off-axis radius and a sector-coded calibration coordinate with nearest-position assignment. The sector code has four quadrant values; it is an interpolation label, not a physical detector azimuth and not a portable PSF descriptor.
Validation uses leave-one-position-out cross-validation: each position is held out, the calibration is built from the remaining positions, and the predicted is compared to the held-out position’s internally calibrated reference limit. This held-out reference limit, , is defined using the same reference fitter that is used throughout Section 4: for each simulated source strength, the median across repetitions is computed, a smooth interpolating spline is fitted to the median versus source-strength curve, and is defined as the exact root where the spline crosses . No information from the held-out position is permitted to leak into the interpolation model, ensuring a strict out-of-sample prediction test. We emphasize that this cross-validation tests the calibration’s ability to self-consistently interpolate the pipeline’s own threshold mapping across the calibration grid; it is not an absolute astrophysical validation against independent observational data or against injection-recovery detection fractions. Such absolute validation, for example comparing the predicted against an injection-recovery 50% completeness flux from dense fake-source injection at representative positions, is a necessary future step for production use. The reported error is
Table [tab:unified-validation] summarizes this conditional interpolation estimand. At , the preferred 2D interpolation yields a median error of 4.1% and a maximum error of 11.0% for positions inside the calibration hull (10 of 14 held-out positions); the nearest-calibration-position assignment covers all 14 positions with a maximum error of 17.2%. Background transfer from to the range – adds at most 5.7% error in at . Both errors decrease at higher thresholds.
lccc Mean crossing: median error (in hull) & 4.1%
& 3.1% & 2.6%
Mean crossing: maximum error (in hull) & 11.0% & 9.1% &
7.8%
Mean crossing: maximum error (all positions) & 17.2% & 12.7%
& 10.1%
Background transfer max
(–)
& 5.7% & 4.1% & 3.2%
The uncalibrated theory has a maximum error of 19.9% at (7.2% median), making it only borderline as a standalone approximation at the most stringent threshold. A global scalar correction fails at (maximum error 21.6%), confirming that the residual is position dependent. The 2D interpolation is preferred inside the calibration hull; nearest-position assignment is the tested fallback for the 14-position sample. These percentages are interpolation errors relative to an internally defined crossing. They are not completeness or absolute sensitivity errors.
The mean-crossing result does not calibrate the standalone width. A historical paired sample of 20,855 rows in 148 cells was analyzed with 14 whole-position folds. An affine relation between reference and standalone describes the mean, but propagating its slope alone underpredicts the width: the held-out log-width RMSE is 0.190244 and the geometric observed-to-predicted ratio is 1.186469. Adding one constant jitter term in quadrature, with median training-fold , improves those values to 0.134238 and 1.040060. Two held-out positions remain substantially miscalibrated. Equation ([eq:eml-jitter]) is therefore a development candidate for that historical data set, not a validated standalone or cross-camera width law.
A larger standalone development test assembled 16,308,000 M1/M2 rows. Its homogeneous primary subset contains M1 band 4, M2 band 1, and M2 band 4, with cells defined in fitted-observable space near theoretical –10. The coefficient-free analytic backbone gives a row-weighted mean residual of and a median width ratio . Exact row-mixture passing-fraction MAE values are 0.06290, 0.10795, and 0.06340 at , 8, and 10. Thus both location and width require a standalone configuration correction.
The frozen PN band-4 reference kernel does not transfer as a combined law. Its positive location correction has the wrong sign: in whole-configuration validation, mean RMSE worsens from 1.21980 for the uncorrected baseline to 2.02904 after transfer. The width component alone improves width-ratio RMSE from 0.21518 to 0.15634, but this partial similarity does not rescue the joint kernel. A newly fitted shared transition reaches mean and width-ratio RMSE of 0.96166 and 0.15273; none of the preregistered injection-PSF feature additions passes both configuration and coordinate gates.
This negative result rejects universal coefficients, not the
location–scale decomposition. The stored injection template is not the
same object as the internal ELLBETA PSF/search surface used by
emldetect, so the test does not show that PSF information
is irrelevant or that a fitter-state-conditioned calibration is
impossible. No multiband PN outcome is used here. A
one-realization-per-cell three-camera execution has established
technical feasibility only and cannot estimate a conditional width.
Once both conditional moments are available, a detection threshold can be treated as a crossing event rather than as a deterministic equality. This section gives two mathematical consequences. A conventional -selected catalog can be interpreted backward in fitted-amplitude space after supplying a prior, or a calibrated kernel can be used forward to define fitted-amplitude percentile crossings. These are consequences of a conditional kernel, not claims that the historical standalone case study or an official XMM catalog already has a validated selection function.
Here denotes the source-amplitude coordinate used by the observable-conditioned theory, consistent with the fitted-amplitude convention used throughout this work. It is not an intrinsic injected flux.
Figure 1 gives the visual version of the construction. A traditional threshold is a horizontal cut in space. Reading the conditional distribution backward from that cut gives the fitted-amplitude distribution of a traditional threshold-selected catalog, once a source-count prior is specified. Reading it forward gives percentile crossings: the usual sensitivity map is the mean crossing, while a lower-percentile crossing defines a more conservative fitted-amplitude threshold-crossing cut.
lll Traditional catalog & Backward:
&
required
Fitted-amplitude cut & Forward:
& Not required
For each detector position, the operational threshold is first converted to the corresponding single-image three-parameter threshold. For example, corresponds to under the convention. The distributional calculation is then carried out in space; remains only the final operational threshold scale.
For a configuration with calibrated location and width functions, define
With the Gaussian residual approximation, the local kernel is summarized as
and the corresponding th percentile curve is
where is the standard-normal quantile. The percentile construction requires, in addition to the calibrated mean residual, a calibrated threshold-band scatter model in space. Thus , , , and .
For the matched PN band-4 reference fitter, and are the functions validated in Section 4. The historical standalone experiment calibrates a mean residual , so its mean can be written within that case study. It does not yet supply a deployable across all positions. Consequently, the lower-percentile formulas are demonstrated for the matched kernel; their standalone use remains conditional on a new width and tail validation.
The first application keeps the traditional catalog definition. A source is selected by an observed detection statistic, for example , equivalently . The new ingredient is that the catalog threshold can now be interpreted through the forward kernel in Equation ([eq:deltac-forward-kernel]).
For a source with a measured value , the fitted-amplitude posterior in catalog space is
where is a source-count prior expressed in the fitted-amplitude coordinate. For a threshold-selected ensemble, the analogous expression is
with
Under the Gaussian residual approximation, this passing probability is a normal survival function evaluated in space. This route does not replace the traditional threshold. Instead, it adds the missing likelihood layer needed to interpret the fitted-amplitude distribution of a -selected catalog. This is also the point at which Eddington bias enters: the noisy threshold statistic must be combined with a source-count prior, especially when the underlying is steep (Eddington 1913; Georgakakis et al. 2008; Wang 2004). A specified value or threshold therefore does not by itself define a flux distribution; it defines a likelihood factor in space. To obtain or , one must combine that likelihood with a source-count prior, for example a power-law slope or index expressed in the fitted-amplitude coordinate. Because this is an inverse problem, it necessarily depends on the population prior. If the desired prior is an intrinsic rather than a prior in fitted catalog amplitude, an additional measurement model connecting to the fitted amplitude is required.
The corresponding fitted-amplitude effective area is probabilistic rather than a step function derived from one limiting amplitude. At a given pixel , detecting a source with fitted amplitude has the catalog-space density
up to normalization over the fitted-amplitude coordinate. Thus the traditional threshold catalog assigns a probability to detected amplitudes at each pixel, rather than a single deterministic limiting amplitude. The catalog-space effective area at fitted amplitude is
where indexes sky or detector pixels. If the goal is the expected number of detected fitted amplitudes in an interval , the same probability-weighted area is combined with the source-count prior,
Equation ([eq:traditional-probabilistic-sky-coverage]) is not a true-flux selection function. The latter requires the joint measurement model that the observable-conditioned kernel deliberately does not supply. For the idealized selection event alone, the true-flux effective area would be Additional candidate-generation and catalog-inclusion rules enlarge the joint state and replace this integration domain by the full selection event. A factorization into is useful only after its conditional-independence and transport assumptions have been checked in matched injection simulations.
This is the sense in which the catalog-space effective area for Application I follows the probability distribution of detected fitted amplitudes, whereas the percentile-cut construction below starts from a chosen fitted-amplitude threshold-crossing probability.
The second application uses the same kernel in the forward direction. Instead of starting from a noisy -selected sample and asking what fitted amplitudes it contains, one asks what fitted amplitude is required for a specified fraction of realizations to exceed the detection threshold. This map-construction step fixes and evaluates , so no slope or source-count index is required.
The usual sensitivity-map calculation is recovered as the mean-curve crossing
Under a symmetric Gaussian residual approximation, this is also the 50th-percentile threshold-crossing point. We avoid calling it an empirical 50% completeness flux, because injection-recovery completeness is usually defined as in the injected-flux coordinate.
More conservative fitted-amplitude cuts are obtained by solving the threshold equation with lower percentile curves,
with . The direction is important. Requiring 90% of realizations at a given fitted amplitude to exceed the threshold uses the lower 10th percentile, not the upper 90th percentile:
Similarly, using the lower envelope corresponds to an approximately 97.7% threshold-crossing criterion under the Gaussian approximation. Table [tab:threshold-crossing-terms] summarizes the terminology used here.
lll Mean limit &
& Corrected fitted-amplitude mean crossing (Asimov when
)
50% threshold crossing &
& Mean crossing if Gaussian
90% threshold crossing &
& 90% expected to pass
Injection 50% completeness &
& Fake-source recovery
Applying the same crossing rule to every detector pixel produces a family of fitted-amplitude sensitivity maps: a mean-threshold map, an 84% threshold-crossing map from the lower 16th percentile, a 90% threshold-crossing map from the lower 10th percentile, and so on. Each map can be converted into a fitted-amplitude effective-area curve by counting the sky area over which the local limiting amplitude is below a trial fitted amplitude,
$$\Omega_{p,\rm fit}(S) = \sum_j \Omega_j\, I\!\left[S_\mathrm{lim}^{(p)}(j) \le S\right],$$
where indexes sky or detector pixels and denotes the percentile curve used in Equation ([eq:percentile-threshold-crossing]). This threshold-crossing effective-area calculation is an instrumental, catalog-space quantity in the same fitted-amplitude coordinate used by the observable-conditioned theory: it depends on the local background, PSF, mask, threshold, and residual calibration, but it does not require an assumed slope. It should not be relabeled as ; Equation ([eq:true-flux-sky-coverage]) is needed for that quantity. A source-count model is needed in subsequent population inference.
Both applications come from the same conditional distribution , but they answer different questions. Application I is backward: it starts from an observed value or a -selected catalog and infers the fitted-amplitude distribution. This is the natural setting for Eddington-bias modeling, because the asymmetry near a detection threshold is produced by combining the noisy threshold statistic with a steep source-count prior.
Application II is forward: it starts from a fitted amplitude and asks for the probability of crossing the threshold. This does not remove population-level Eddington bias from a luminosity-function inference, but it allows one to construct an instrumental catalog-space subsample with an explicitly stated threshold-crossing probability. In this sense, the percentile sensitivity map is a catalog-construction tool, whereas the threshold-selected posterior is a catalog-interpretation tool.
True-flux deboosting is a larger problem. It requires , a prior on , and normalization by the same selection event used to form the catalog. The conditional kernel is one factor in that calculation, not a complete Eddington correction by itself.
Section 4 validates threshold-passing probabilities for the matched reference fitter. Section 5 measures only the interpolation error of the historical standalone mean crossing: at , the maximum is 11.0% inside the calibration hull and 17.2% for nearest-position assignment across all 14 held-out positions. The standalone width and tails are not closed. Before a lower-percentile standalone map is used, predicted passing probabilities must be compared with empirical passing fractions in new whole-position simulations. An official catalog adds candidate-generation and inclusion rules that require a separate end-to-end selection test.
Thus the application presented here is a distributional framework for threshold interpretation and fitted-amplitude threshold crossing. Full injection-recovery simulations remain the validation standard for a true-flux selection function in complex survey fields.
The calibration grid, cross-validation tables, figure-generation scripts, configuration files, and minimal scripts reproducing key validation results will be made available upon publication via a persistent Zenodo archive.
The evidence tiers are summarized in Table [tab:evidence-tiers]. The analytical decomposition is the portable result. The numerical and functions are attached to a particular effective PSF, mask, search domain, and fitter. Changing camera or energy band changes the PSF, but changing the pipeline can also change the internal PSF raster and optimization state even when the stored injection template is held fixed. Quantitative transfer therefore requires matched calibration rather than a camera label alone.
lll Analytic
& Derived at local Fisher order & Matched three-parameter fits
in the validated likelihood regime
PN band-4 reference
& OOS validated & Reference-fitter threshold probabilities in
tested support
PN band-4 standalone mean grid & LOPO case study & Internal
mean-crossing interpolation only
M1/M2 standalone transfer & Unchanged PN coefficients falsified on
development data & Recalibrate
PN multiband standalone & Not scientifically opened & No
coefficient, width, or tail claim
Official 4XMM catalogs & Not validated here & No weak-source
completeness or catalog correction
The matched simulations are single-image, isolated point-source fits on a locally specified background. Extended sources, crowded fields, multi-image likelihoods, merged observations, background-estimation uncertainty, and a different number of fitted degrees of freedom are outside the derivation as tested. The reference kernel is validated in PN band 4. The historical standalone mean grid covers – only; its sector coordinate is not a physical PSF parameter. The M1/M2 experiment is a strong negative transfer test, but the exact internally rendered ELLBETA fit/search object was not saved, so it cannot determine which missing fitter-state variable would restore portability.
No catalog-selected sample is used as core evidence for the weak-source kernel. Near a catalog threshold, the unselected parent distribution is unknown; a likelihood-selected catalog cannot by itself identify the conditional width of the missing sources. Official catalog selection also includes stacking, confusion, vignetting, background-map variation, and joint multi-band decisions. End-to-end simulations used by other survey pipelines illustrate why these effects are normally calibrated as a complete system (Brunner et al. 2022; Evans et al. 2024; Liu et al. 2022).
Finally, all Section 6 effective-area curves before Equation ([eq:true-flux-sky-coverage]) live in the fitted-amplitude coordinate. Population-level Eddington correction and true-flux sky coverage require a source prior and a joint measurement model connecting , , and .
The main analytical result is the leading Fisher approximation to the conditional width The Fisher projection removes the amplitude and two position score directions absorbed by the fit. It is evaluated from fitted observables, the pixelized PSF, background, and mask. The companion expression extends the amplitude-only analytical Cash-sensitivity calculation of Stewart (2009); Stewart (2006) provides earlier matched-filter sensitivity context. The variance is the element that turns a deterministic threshold curve into a distribution, with supplying the empirically measured low-count width correction.
In the matched PN band-4 reference fitter, 266,837 high-signal fits give a standardized residual mean of 0.068 and width of 0.977. At low signal, the empirical location–scale kernel describes 810,987 clean fits. In realization-block validation, the mean passing-fraction error over , 8, and 10 decreases from 0.013585 for raw Fisher theory to 0.001999 after both corrections, close to the 0.001967 empirical-bin benchmark. The width improvement is smaller than the location improvement but reproducible. A grouped test also rejects replacing by a function inferred from alone. The mean centers the residual and removes cell gradients; it does not determine the conditional width.
Standalone emldetect remains scientifically useful in
this framework, but as a pipeline-specific layer. The historical PN
band-4 experiment quantifies interpolation of an internally defined mean
crossing. Its 11–17% errors are not completeness errors, and its width
is not closed. The M1/M2 development experiment provides the
complementary boundary: unchanged PN coefficients fail to transfer. In
the primary
theoretical-–10
development cells, the row-weighted location residual is about
and the median observed-to-Fisher width ratio is about 0.84. The
reusable result is therefore the analytical location–scale architecture,
not a universal coefficient table.
The calibrated matched kernel defines fitted-amplitude threshold-crossing probabilities and a fitted-amplitude effective area. Backward catalog interpretation requires a prior. True-flux completeness and Eddington correction additionally require the joint measurement model and the catalog’s full selection normalization. These boundaries preserve the practical value of the new form while keeping its empirical coefficients tied to the PSF/search/fitter configuration in which they were measured.
The calibration grid, cross-validation tables, figure-generation
scripts, SAS configuration files, and minimal scripts reproducing key
validation results (including the exact commands, random seeds, PSF
parameters, and emldetect invocation flags) will be
deposited in a persistent archive upon publication. The DOI and license
will be inserted before submission. Machine-readable calibration tables
and the code used to generate the figures are maintained with the
project artifacts during review.
R.H. developed the statistical model, simulations, analysis code, and initial manuscript. Draft note: coauthor roles will be completed using the CRediT taxonomy after all authors approve the submission version.
Draft note: the competing-interest declaration will be confirmed by all authors before submission.
This work acknowledges support from the National Natural Science Foundation of China. Draft note: grant identifiers will be added after author confirmation.
This study uses numerical simulations and archival calibration products; it involves no human participants, animals, or personally identifiable data.
Draft note: a disclosure consistent with the policy of the selected journal will be finalized before submission.
Figure 2 is a diagnostic check on the conditioning variable. The simulations know the injected source strength , but the theory and the matched calibration condition on the fitted amplitude . At fixed fitted , the standardized residual distributions from different subsets overlap closely. The residual offset in the lowest bin is the same low-signal beyond-Fisher bias discussed in Section 4.2, not evidence for conditioning on .
This research has made use of data obtained from the XMM-Newton satellite, an ESA science mission with instruments and contributions directly funded by ESA Member States and NASA.