The engine every Ardia model thinks with. Three layers — deterministic policy, clinical reasoning, and a denial-pattern library — so an answer is grounded in the rule it came from, not guessed. Auditable end to end.
Each layer does one job, and the boundary between deterministic rules and language-model reasoning is deliberate — it is what keeps TARA auditable.
Deterministic encoding of payer LCD/NCD and MolDX policy, frequency and panel caps, and bundling logic. Rules are versioned and cited — never hallucinated. When TARA says a claim fails a policy, it can point at the exact clause.
A frontier language model (Claude), behind guardrails, reads the de-identified record against the matched policy and finds the medical-necessity gap — then drafts the argument with citations. The model reasons; it does not get to override the policy layer or the safety rules.
Payer-specific CARC/RARC playbooks — the appeal angle for each denial code, learned from real denial patterns. It is also what lets TARA decline a correct denial rather than paper over it.
A denied toxicology claim, reasoned to a recovery narrative. Illustrative, synthetic claim — production runs the full MAC-specific policy set under a BAA.
TARA is honestly not a trillion-parameter foundation model, and we will not pretend it is. It is an architecture: a deterministic policy engine wrapped around a frontier language model, grounded in retrieval, and gated by governance. That design is the point.
Every clinical or policy claim is grounded in a retrieved, cited source — payer LCDs, NCCN, CPIC — so the model argues from the record, not from memory.
The rule layer is code, not prose the model can talk its way around. A frequency cap is a frequency cap.
PHI is de-identified before the model sees it; every access is written to a hash-chained audit trail; access is role-scoped and minimum-necessary.
Nothing ships without passing the safety and regression suite. A broken guardrail blocks the deploy.
A reasoning engine can't honestly be scored like a classifier on one accuracy number. Here is what is measured today, stated plainly.
The de-identification + audit kernel (PHI de-identification to HIPAA Safe Harbor, hash-chained audit, RBAC / minimum-necessary) is being built. Its test suite is a modelled target — not yet measured.
The release-gating harness (refusal, escalation, de-identification, regression) is being built — generalising the subject-independent protocol that produced Cadence's numbers. A modelled target — not yet measured.
Measured today by policy-match correctness on synthetic cases. End-to-end appeal accuracy will be measured on a design-partner dataset under a BAA before any performance claim is published.
No appeal win-rate, denial-recovery percentage, or clinical-outcome figure is claimed for TARA, because Ardia is pre-revenue and those numbers cannot honestly be known before a pilot. The one genuinely measured model is Cadence (see the Test Results page); Sentinel and Crucible above are honest works-in-progress. Everything else here describes the architecture and its intended behaviour, with synthetic examples labelled as such. A published benchmark is on the roadmap — built the same way Cadence's was: honestly, and measured once.
See the full model lineup, or the trained model with real benchmarks — Cadence.