Validated on synthetic recovery. Real-world validation begins only when a consented, sufficiently connected outcome cohort is available.
This distinction is what a serious academic partner expects, and it does not diminish the work. Everything below labelled "future" or "synthetic" is exactly that.
Identical conditions, different people
An aggregate suitability score conflates two different things: how hard a condition objectively is, and the skill distribution of the population that historically shows up there. The same eighteen knots of wind is a routine day for one rider and the edge of control for another.
Because those two effects are mixed, two spots with identical population-suitability curves can sit on very different underlying difficulty, masked by different clienteles. The method proposed here tries to pull those apart: an estimated condition-difficulty function, separate from a latent per-person skill.
An extension of the suitability score, not a new engine
An outcome is a triple: rider r, condition (metric value x at sub-spot s), and a behavioural outcome y. Classic Item Response Theory (IRT) models discrete test items; the proposed adaptation treats a continuous physical condition as the axis, so difficulty becomes a smooth function rather than a per-question value.
Integrating skill out under the population distribution recovers exactly today's single suitability curve, so this is strictly an extension of the existing score. When the rider by condition graph is not connected, the method does not force a fit: it falls back to the explicit single curve. Preferring an explicit regression to a non-identified fit is the intended default.
P(y=1 | r, x, s) = σ(a · (θ_r − δ(x, s)))
Latent user skill. One model-specific scalar per pseudonym, identified by an N(0,1) location and scale.
Estimated condition difficulty. A model-estimated latent conditioned on the included variables, the user sample, the outcome definition, the physics prior, graph connectivity and missingness. "Intrinsic difficulty" is a hypothesis about this quantity, not an observed fact.
Discrimination. How sharply the condition separates skill levels.
What it can estimate, what it cannot prove yet
Given a connected, consented cohort that passes the diagnostics.
- A latent skill scalar per pseudonym, up to the model's identification.
- An estimated condition-difficulty function δ(x, s), conditioned on the included variables and the outcome definition.
- A skill-conditioned suitability curve that reduces to the population curve when skill is integrated out.
- A held-out predictive comparison against the single-curve baseline.
The only evidence so far is synthetic recovery.
- Any real-world accuracy claim. No real outcome cohort has been fitted.
- That δ is the true intrinsic difficulty. Intrinsic difficulty is a model interpretation, not an observed fact.
- Identification on every spot. Many cells will never clear the connectivity and coverage bar and stay on the single curve.
- A certification of anyone's ability, or any eligibility, pricing or safety use.
From physics prior to a gated recommendation
Physics profile
A versioned expert curve enters as the prior. It anchors difficulty wherever outcome data is sparse.
Consented outcomes
Paired (condition, outcome) records from a consented cohort supply the behavioural evidence.
Eligibility and diagnostics
Coverage, sample size, connectivity and out-of-sample checks decide whether a fit is served at all.
Estimated difficulty
δ(x, s) is estimated, conditioned on the data, the prior and the outcome definition.
Skill-conditioned recommendation
Suitability is refined per skill level, and never below the hard safety gates.
The IRT layer never relaxes physical hard gates. It only refines suitability among conditions that have already passed them.
Never overrides safety gates
The prior refines suitability only among conditions that already pass the hard physical gates. It cannot reopen a gated condition.
Weight scales with evidence
The prior dominates where outcomes are sparse and yields to the data as a cell accumulates observations.
Varies by region and spot
The physics profile is defined per activity and can differ by region and sub-spot.
Versioned
Each profile carries a version, and every fit pins the version it used.
Updated only after review
Outcomes can revise a profile, but only through a reviewed, versioned release, never silently.
No retroactive rewrite
A new profile version does not rewrite historical scores. Past fits stay pinned to the version they used.
Defining the target properly
"Good session or not" is not a clean research target. The intended outcome taxonomy separates non-observation, cancellation and operator judgement, and governs safety events apart from everything else.
- 01Not observed
- 02Cancelled for weather
- 03Cancelled for a non-weather reason
- 04Ran, outcome unknown
- 05Ran, operator-rated unsuitable
- 06Ran, acceptable
- 07Ran, favourable
- 08Safety incident or near miss (separately governed, never inferred)
Selection bias is the core operational risk. Skilled people choose harder conditions, and operators cancel when they anticipate problems, so a "not" label is not a clean failure. The validation strategy has to model this, not assume it away.
Identification, reproducibility and privacy
Identification
Connectivity of the rider by condition graph is necessary, but it is only the first eligibility check. The model is served only when minimum coverage, sample-size and out-of-sample diagnostics also pass.
Reproducibility
Every fit is pinned to a cohort hash, model version and reproducible method specification. Authorised reviewers can reproduce the published fit from the documented inputs and release artifact.
Privacy
Individual-level posterior data is deleted upon a valid request. Released atlas aggregates are retained only after documented de-identification and disclosure-control review.
Connectivity is the first eligibility check. A fit is served only when every check below also passes.
- Minimum riders
- Minimum outcomes per rider
- Condition overlap between riders
- Coverage of the condition range
- Maximum missingness
- Minimum effective sample size per spot and regime
- Posterior-uncertainty diagnostics
- Holdout calibration and performance
- Sensitivity to priors and outcome definition
A difficulty estimate from a small cohort, a rare spot or distinguishable users can retain re-identification risk, or be a derived datum linked to a person. GDPR Recital 26 separates pseudonymisation from anonymisation and requires assessing the means reasonably likely to be used, so an atlas aggregate is not treated as automatically anonymous.
Methodology note: estimation uses Marginal Maximum Likelihood with Gauss-Hermite quadrature, so a fixed seed yields the same fit. Full Bayes is documented as a future upgrade for richer uncertainty.
Recovery on synthetic data only
The synthetic-recovery proof generates a cohort with known θ per rider and known δ(x), fits the model, and checks whether it recovers the ground truth. It runs in under a minute on a laptop and touches no real data. It is the artifact reviewers can inspect, and it demonstrates the estimator works where the truth is known, nothing more.
On the reference cohort of 80 riders by 30 outcomes, the fit recovers θ at correlation 0.96 against truth, locates the difficulty minimum within 3 units of ground truth, and improves the held-out Brier Skill Score by 0.33 over the expert-curve baseline. These are synthetic-recovery figures, not real-world accuracy.
$ uv run inverse-suitability latent-demo --out /tmp/demo
=== Inverse Suitability: synthetic two-factor cohort ===
{
"theta_recovery_corr": 0.96,
"difficulty_min_knots": 16.7,
"discrimination_a": 2.16,
"marginal_matches": true,
"held_out_bss": 0.33,
"identification_ok": true,
"n_train": 2400
}Real-world validation protocol
Before any deployment or difficulty endpoint, the method is put through a preregistered validation. The protocol is designed to be falsifiable and to fail cleanly.
- 01Preregistered hypotheses
- 02Initial population and activity
- 03Outcome definition
- 04Inclusion and exclusion criteria
- 05Selection-bias strategy
- 06Temporal train / validation / test split
- 07Baselines to beat
- 08Calibration, discrimination and fairness metrics
- 09Go / no-go criteria
- 10Governance, ethics and consent
- 11A non-deployment policy if it fails
Kitesurfing, one region, one outcome taxonomy, one defined cohort protocol. Deliberately narrow: more credible, and more fundable, than many activities and spots at once.
Skill is a model-specific latent estimate, not a certification of ability. It is not used for eligibility, pricing or safety clearance. Estimating skill from outcomes can absorb access to coaching, gear, free time, economic means and operator bias, even without collecting protected attributes.
A research track, not a shipped feature
Skill-conditioned scoring and the difficulty endpoint are not current Pro+ features. They depend on real outcomes and a connected graph, neither of which exists today. Here is the honest state of each element.
| Element | Status |
|---|---|
| Paper and synthetic code | Available now |
| Method and protocol | Research track |
| Outcome data collection | Design-partner phase |
| Real-world validation | Not yet complete |
| Skill-conditioned production scoring | Future, subject to validation |
| Difficulty endpoint | Future, pilot-only after eligibility checks |
Potential uses, all conditional
Personalisation
Skill-conditioned recommendation for a returning rider. Future, subject to validation.
Difficulty atlas
A research artifact describing estimated condition difficulty per spot. Future, pilot-only after eligibility checks.
Funding posture
The methodology may provide a research work-package or proposal foundation for relevant Horizon Europe, Interreg and environmental-observation calls.
Academic collaboration
Seeking academic collaborators for real-world validation, study design and independent evaluation.
Investment narrative
If validated on real outcomes, the methodology could create a differentiated, longitudinal evidence layer beyond environmental measurements alone.
Risk analysis
At most, exploratory risk analysis. Insurance pricing needs empirical validation, stability, causal analysis, regulatory review and partner approval; a synthetic proof implies no actuarial readiness.
Seeking academic collaborators for validation
"Inverse Suitability: Identifying Condition Difficulty and Rider Skill from Behavioural Outcomes via Continuous-Item Response Theory."
Carucci, F. (2026). Inverse Suitability: Identifying Condition Difficulty and Rider Skill from Behavioural Outcomes via Continuous-Item Response Theory. arXiv:2607.01961 [stat.AP].
We are seeking academic collaborators for real-world validation, study design and independent evaluation. A serious collaboration defines these up front: