orthoharness

An open benchmark harness for orthopaedic AI

The models are good enough. The test is not.

A harness is everything around the model: what goes in, which tools it can use, what checks the output, how it is scored, and what gets recorded. orthoharness scores orthopaedic AI, and surgeons, the way surgeons actually decide. Apache-2.0, no dependencies, runs offline.

Status: a reference implementation, not a validated benchmark. Its six demo cases are synthetic, written to exercise the harness, and are not clinical guidance. Its loss weights are placeholders that show the shape. The real cases and weights belong to the specialty.

Why a harness

Orthopaedics does not need its own model. It needs its own harness.

General-purpose models now beat specialized clinical tools on benchmarks (Vishwanath, Nature Medicine 2026). What decides whether one is safe in an orthopaedic clinic is everything around it, and a scorecard built for how surgeons actually decide.

Start with what goes in. Asked for evidence with nothing to read from, a model fabricated 7.9% of 2,556 citations (McCavitt, JBJS OA 2026). Handed the source papers, models hallucinated in 0.08–6% of extractions, and leaving things out became the main error, 60–74% of the total (Shankar, J Biomed Inform 2026). Different tasks and models, so read it as direction, not rate: grounding the input removes most of the invention and moves the error to what is missing. A harness has to measure both.

The nine rows

What today's tests miss, and what the harness does instead.

  1. 1

    Fixed questions, asked once.

    The record is revealed one stage at a time. Committing before the record supports it is scored as premature.

    Gong · JMIR 2025: 39 benchmarks, practice far below knowledge. AgentClinic · npj Digital Medicine 2026.

  2. 2

    One yardstick, one score.

    Two yardsticks, the guideline and the treating attending, scored apart and never averaged. Where they disagree, it shows which side the system took.

    Dagher · CORR 2024: 90 of 100 with the guideline, 78 with the attending.

  3. 3

    Guessing pays.

    “Not enough information” is a real answer with its own row in the loss matrix. A correct “do not operate” scores the same as a correct “operate”.

    CliniCARE-Bench · preprint: 16 of 16 systems over-committed.

  4. 4

    Only the answer is graded.

    The route is graded: each case lists the contraindication checks a safe answer must show.

    PLOS Medicine 2026: advice without a contraindication check is a process-safety failure.

  5. 5

    A citation passes if it exists.

    Every DOI is looked up at CrossRef and its title compared with the one given. Whether the paper supports the claim is counted as unread until a person reads it.

    McCavitt · JBJS OA 2026: 7.9% fabricated; about nine in ten real ones misattributed.

  6. 6

    Only invented content counts.

    Omissions count: each case lists the facts the reasoning has to name.

    Shankar · J Biomed Inform 2026: with the source supplied, omissions were 60–74% of errors.

  7. 7

    Only the machine is measured.

    The same grader scores a surgeon's own answers, so aided and unaided can be compared on the same cases.

    Budzyń · Lancet Gastroenterol Hepatol 2025: adenoma detection 28.4% before AI exposure, 22.4% after.

  8. 8

    Written once, then stale.

    The keep-or-strike delta turns a draft and a signed note into labeled kept, struck and added lines.

    The evidence a person looked, and new data from real practice.

  9. 9

    Nobody owns the answer.

    Every run records who signed. Without a signer, answers are reported as unsigned drafts. There is no single overall score.

    California AB 1979, signed Sep 30, 2026.

Publish the loss matrix and weight each error by the harm it does: a missed infection is not an extra X-ray. Report results by kind of error, not as one leaderboard number. Withhold the test questions so no model trains on the exam.

Run it

Clone it, run the tests, score something.

git clone https://github.com/blainomd/orthoharness && cd orthoharness
npm test        # 13 checks, offline

node bin/orthoharness.mjs run cases/demo --adapter always-operate --offline
node bin/orthoharness.mjs run cases/demo --adapter replay --answers examples/answers-careful.jsonl

Any OpenAI-compatible endpoint, hosted or local. The key is read from the environment and never written to a run file:

ORTHOHARNESS_BASE_URL=http://localhost:11434/v1 ORTHOHARNESS_MODEL=llama3.1 \
  node bin/orthoharness.mjs run cases/demo --adapter chat

Score a person

The replay adapter reads a surgeon's own recorded answers and grades them by the same rules. compare puts aided and unaided side by side.

Keep the delta

delta draft.txt signed.txt labels every kept, struck and added line. A note signed exactly as drafted is flagged: no evidence anyone edited it.

Keep the exam private

hash gives a sha256 commitment to a case set. Publish the hash and the loss matrix; withhold the questions.

Before you buy

Six questions for any AI vendor. Including us.

Scribes, coders, prior-auth and appeal agents are all selling to orthopaedic practices. Most publish an outcome they promise, not an error rate they measured. The harness, turned into questions:

1 What is your error rate, against what reference, on which cases?

A number measured against a named reference on a stated case mix. Knowledge scores run far above practice scores across 39 benchmarks (Gong, JMIR 2025), so ask for practice.

2 Can it say “not enough information”, and how often does it?

A system that never abstains is guessing on the incomplete records. All 16 systems scored with abstention over-committed (CliniCARE-Bench, preprint).

3 What share of drafts do surgeons change before signing?

“Signed without edits” is not a feature. It means there is no evidence anyone looked. Drafted histories carried erroneous information 36% of the time in a JAAOS trial.

4 Does every citation in your letters say what the letter claims?

Existing is not supporting: 7.9% fabricated and about nine in ten real citations misattributed (McCavitt, JBJS OA 2026). Ask who reads them before they go to a payer.

5 Do you sell defensibility or uplift?

A promised rise in code levels is what an audit looks for. Ask what happens to a claim when the documentation does not support it: does the tool say no?

6 Will you run it through an open harness, and publish the result?

Two yardsticks, staged cases, credit for abstaining, the delta, a named signer. This harness is open source; any vendor can run it.

Orthopaedics already ran this test on its last big technology. In 2026, three studies measured robotic knee replacement instead of trusting the brochure:

A RASKAL, a randomized trial

Robotic-assisted surgery was not superior to computer-assisted surgery, and functional alignment was not superior to mechanical alignment, on KOOS-12 up to two years. MacDessi · Bone Joint J 2026.

B RACER-Knee, a randomized trial

Robotic-arm-assisted knee replacement gave the same patient outcomes as conventional surgery, with longer operations at higher cost. Parsons · Lancet 2026.

C The UK registry, a target trial emulation

Five-year implant survival after total knee replacement was 98.5% conventional and 98.6% robotic, with no difference in revision risk (HR 1.03). National Joint Registry data · BMJ 2026.

Nobody had to take a vendor's word for the robot. The same three questions apply to AI: a randomized comparison, a real-world registry, and a cost on the record.

Who should own it

Not a vendor. Including us.

The cases, the loss weights and the held-out test set belong with the specialty's own societies, journals and registries, versioned and independent. This is the plumbing: fork it, replace the demo cases, publish your loss matrix and your hash.

Contributions are welcome: synthetic cases, or properly de-identified ones with the right approvals (never patient information in an issue or pull request), graders that replace phrase matching, and adapters.

A promised outcome is marketing. A measured error is evidence.