UTC --:--:--
The Sustainability Register · Research

A continuous-item IRT formulation for outdoor suitability

A research method that separates the estimated difficulty of a condition from the latent skill of the person in it, fitted from behavioural outcomes. It is a proposed extension of the ordinary suitability score, not a replacement engine, and it is a method, not a shipped product.

Current status

Validated on synthetic recovery. Real-world validation begins only when a consented, sufficiently connected outcome cohort is available.

This distinction is what a serious academic partner expects, and it does not diminish the work. Everything below labelled "future" or "synthetic" is exactly that.

01
The problem

Identical conditions, different people

An aggregate suitability score conflates two different things: how hard a condition objectively is, and the skill distribution of the population that historically shows up there. The same eighteen knots of wind is a routine day for one rider and the edge of control for another.

Because those two effects are mixed, two spots with identical population-suitability curves can sit on very different underlying difficulty, masked by different clienteles. The method proposed here tries to pull those apart: an estimated condition-difficulty function, separate from a latent per-person skill.

02
The formulation

An extension of the suitability score, not a new engine

An outcome is a triple: rider r, condition (metric value x at sub-spot s), and a behavioural outcome y. Classic Item Response Theory (IRT) models discrete test items; the proposed adaptation treats a continuous physical condition as the axis, so difficulty becomes a smooth function rather than a per-question value.

Integrating skill out under the population distribution recovers exactly today's single suitability curve, so this is strictly an extension of the existing score. When the rider by condition graph is not connected, the method does not force a fit: it falls back to the explicit single curve. Preferring an explicit regression to a non-identified fit is the intended default.

The model

P(y=1 | r, x, s) = σ(a · (θ_r − δ(x, s)))

θ_r

Latent user skill. One model-specific scalar per pseudonym, identified by an N(0,1) location and scale.

δ(x, s)

Estimated condition difficulty. A model-estimated latent conditioned on the included variables, the user sample, the outcome definition, the physics prior, graph connectivity and missingness. "Intrinsic difficulty" is a hypothesis about this quantity, not an observed fact.

a

Discrimination. How sharply the condition separates skill levels.

03
Scope

What it can estimate, what it cannot prove yet

Can estimate

Given a connected, consented cohort that passes the diagnostics.

  • A latent skill scalar per pseudonym, up to the model's identification.
  • An estimated condition-difficulty function δ(x, s), conditioned on the included variables and the outcome definition.
  • A skill-conditioned suitability curve that reduces to the population curve when skill is integrated out.
  • A held-out predictive comparison against the single-curve baseline.
Cannot prove today

The only evidence so far is synthetic recovery.

  • Any real-world accuracy claim. No real outcome cohort has been fitted.
  • That δ is the true intrinsic difficulty. Intrinsic difficulty is a model interpretation, not an observed fact.
  • Identification on every spot. Many cells will never clear the connectivity and coverage bar and stay on the single curve.
  • A certification of anyone's ability, or any eligibility, pricing or safety use.
04
The pipeline

From physics prior to a gated recommendation

1

Physics profile

A versioned expert curve enters as the prior. It anchors difficulty wherever outcome data is sparse.

2

Consented outcomes

Paired (condition, outcome) records from a consented cohort supply the behavioural evidence.

3

Eligibility and diagnostics

Coverage, sample size, connectivity and out-of-sample checks decide whether a fit is served at all.

4

Estimated difficulty

δ(x, s) is estimated, conditioned on the data, the prior and the outcome definition.

5

Skill-conditioned recommendation

Suitability is refined per skill level, and never below the hard safety gates.

Safety invariant

The IRT layer never relaxes physical hard gates. It only refines suitability among conditions that have already passed them.

The physics prior, in detail

Never overrides safety gates

The prior refines suitability only among conditions that already pass the hard physical gates. It cannot reopen a gated condition.

Weight scales with evidence

The prior dominates where outcomes are sparse and yields to the data as a cell accumulates observations.

Varies by region and spot

The physics profile is defined per activity and can differ by region and sub-spot.

Versioned

Each profile carries a version, and every fit pins the version it used.

Updated only after review

Outcomes can revise a profile, but only through a reviewed, versioned release, never silently.

No retroactive rewrite

A new profile version does not rewrite historical scores. Past fits stay pinned to the version they used.

05
The outcome

Defining the target properly

"Good session or not" is not a clean research target. The intended outcome taxonomy separates non-observation, cancellation and operator judgement, and governs safety events apart from everything else.

  1. 01Not observed
  2. 02Cancelled for weather
  3. 03Cancelled for a non-weather reason
  4. 04Ran, outcome unknown
  5. 05Ran, operator-rated unsuitable
  6. 06Ran, acceptable
  7. 07Ran, favourable
  8. 08Safety incident or near miss (separately governed, never inferred)

Selection bias is the core operational risk. Skilled people choose harder conditions, and operators cancel when they anticipate problems, so a "not" label is not a clean failure. The validation strategy has to model this, not assume it away.

06
Governance

Identification, reproducibility and privacy

Identification

Connectivity of the rider by condition graph is necessary, but it is only the first eligibility check. The model is served only when minimum coverage, sample-size and out-of-sample diagnostics also pass.

Reproducibility

Every fit is pinned to a cohort hash, model version and reproducible method specification. Authorised reviewers can reproduce the published fit from the documented inputs and release artifact.

Privacy

Individual-level posterior data is deleted upon a valid request. Released atlas aggregates are retained only after documented de-identification and disclosure-control review.

Eligibility checklist

Connectivity is the first eligibility check. A fit is served only when every check below also passes.

  • Minimum riders
  • Minimum outcomes per rider
  • Condition overlap between riders
  • Coverage of the condition range
  • Maximum missingness
  • Minimum effective sample size per spot and regime
  • Posterior-uncertainty diagnostics
  • Holdout calibration and performance
  • Sensitivity to priors and outcome definition
Disclosure control

A difficulty estimate from a small cohort, a rare spot or distinguishable users can retain re-identification risk, or be a derived datum linked to a person. GDPR Recital 26 separates pseudonymisation from anonymisation and requires assessing the means reasonably likely to be used, so an atlas aggregate is not treated as automatically anonymous.

Methodology note: estimation uses Marginal Maximum Likelihood with Gauss-Hermite quadrature, so a fixed seed yields the same fit. Full Bayes is documented as a future upgrade for richer uncertainty.

07
Synthetic evidence

Recovery on synthetic data only

Every number in this section is from a synthetic cohort. No real-world validation has been run.

The synthetic-recovery proof generates a cohort with known θ per rider and known δ(x), fits the model, and checks whether it recovers the ground truth. It runs in under a minute on a laptop and touches no real data. It is the artifact reviewers can inspect, and it demonstrates the estimator works where the truth is known, nothing more.

On the reference cohort of 80 riders by 30 outcomes, the fit recovers θ at correlation 0.96 against truth, locates the difficulty minimum within 3 units of ground truth, and improves the held-out Brier Skill Score by 0.33 over the expert-curve baseline. These are synthetic-recovery figures, not real-world accuracy.

Reproducibility · the CLI
$ uv run inverse-suitability latent-demo --out /tmp/demo

=== Inverse Suitability: synthetic two-factor cohort ===

{
  "theta_recovery_corr":   0.96,
  "difficulty_min_knots":  16.7,
  "discrimination_a":      2.16,
  "marginal_matches":      true,
  "held_out_bss":          0.33,
  "identification_ok":     true,
  "n_train":               2400
}
512182532SUITABILITY · DIFFICULTYWIND SPEED (kn)δ MIN · 18 kn
δ(x, s) · estimated difficultyσ(a·(θ−δ)) · expert (θ=+1)σ(a·(θ−δ)) · intermediate (θ=0)σ(a·(θ−δ)) · beginner (θ=−1)
Illustrative synthetic example. The dashed line is an estimated difficulty δ(x, s); the three solid curves are the suitability profiles it implies for three skill levels. It shows the intended separation, not a measured result.
08
The plan

Real-world validation protocol

Before any deployment or difficulty endpoint, the method is put through a preregistered validation. The protocol is designed to be falsifiable and to fail cleanly.

  1. 01Preregistered hypotheses
  2. 02Initial population and activity
  3. 03Outcome definition
  4. 04Inclusion and exclusion criteria
  5. 05Selection-bias strategy
  6. 06Temporal train / validation / test split
  7. 07Baselines to beat
  8. 08Calibration, discrimination and fairness metrics
  9. 09Go / no-go criteria
  10. 10Governance, ethics and consent
  11. 11A non-deployment policy if it fails
Initial validation scope

Kitesurfing, one region, one outcome taxonomy, one defined cohort protocol. Deliberately narrow: more credible, and more fundable, than many activities and spots at once.

Fairness note

Skill is a model-specific latent estimate, not a certification of ability. It is not used for eligibility, pricing or safety clearance. Estimating skill from outcomes can absorb access to coaching, gear, free time, economic means and operator bias, even without collecting protected attributes.

09
Status

A research track, not a shipped feature

Skill-conditioned scoring and the difficulty endpoint are not current Pro+ features. They depend on real outcomes and a connected graph, neither of which exists today. Here is the honest state of each element.

ElementStatus
Paper and synthetic codeAvailable now
Method and protocolResearch track
Outcome data collectionDesign-partner phase
Real-world validationNot yet complete
Skill-conditioned production scoringFuture, subject to validation
Difficulty endpointFuture, pilot-only after eligibility checks
10
After validation

Potential uses, all conditional

Personalisation

Skill-conditioned recommendation for a returning rider. Future, subject to validation.

Difficulty atlas

A research artifact describing estimated condition difficulty per spot. Future, pilot-only after eligibility checks.

Funding posture

The methodology may provide a research work-package or proposal foundation for relevant Horizon Europe, Interreg and environmental-observation calls.

Academic collaboration

Seeking academic collaborators for real-world validation, study design and independent evaluation.

Investment narrative

If validated on real outcomes, the methodology could create a differentiated, longitudinal evidence layer beyond environmental measurements alone.

Risk analysis

At most, exploratory risk analysis. Insurance pricing needs empirical validation, stability, causal analysis, regulatory review and partner approval; a synthetic proof implies no actuarial readiness.

11
Publication and collaboration

Seeking academic collaborators for validation

Published paper

"Inverse Suitability: Identifying Condition Difficulty and Rider Skill from Behavioural Outcomes via Continuous-Item Response Theory."

Carucci, F. (2026). Inverse Suitability: Identifying Condition Difficulty and Rider Skill from Behavioural Outcomes via Continuous-Item Response Theory. arXiv:2607.01961 [stat.AP].

We are seeking academic collaborators for real-world validation, study design and independent evaluation. A serious collaboration defines these up front:

Research questionIP and ownershipData accessEthicsAuthorship criteriaPreregistrationFundingRoles and independence