ToxIQ is a toxicology go-to-market label on MolecuIQ's engine path — the browser rewrites it to molec before the call — that reasons over drug-testing coverage, coding and medical necessity against a small hand-curated set of CMS urine-drug-testing LCDs and produces a cited draft appeal whose quality has never been graded by a coder, biller, compliance officer or payer; a live demo with no toxicology-specific model, no NCCI/MUE or CLFS data, no claim ingestion, and zero customers, pilots, BAAs, revenue or measured accuracy.
Pillar 2 — molecular and genomic diagnostics, read as lab reimbursement: toxicology is the sibling SKU to MolecuIQ and shares its engine path outright. Direct exposure to Pillar 3, the PAMA CLFS repricing: every dollar ToxIQ is intended to defend would be a CLFS dollar exposed to the cut Meridian is intended to model. ToxIQ has defended $0 to date. Meridian's deterministic CLFS/PAMA arithmetic does exist and is unit-tested (models/meridian/clfs.py, mirrored by a client-side calculator on model-pama.html), but it is not wired into the Studio answer path — a Meridian question in the Studio is answered by Gemini reasoning as TARA, so a buyer meets two different Meridians. Pillar 1, precision medicine, touches ToxIQ only through the CPIC/PGx metabolizer overlay the site advertises and the codebase does not contain. There is no clinical or commercial connection to Pillar 4 (pulmonary — PulmoIQ, asthma/COPD) or Pillar 5 (elder care — Aria, Ardia One, Cadence), and this document will not manufacture one. The single real link across all five pillars is architectural: ToxIQ, MolecuIQ, PulmoIQ, Meridian, Aria and Lumen are framings over four engine paths and one six-gate Crucible, so a prompt regression or a silent model roll in that shared substrate degrades products across every pillar at once. That is the honest pillar story: ToxIQ is Pillar 2 commercially, Pillar 3 economically, and coupled to Pillars 1, 4 and 5 only by a substrate with no evaluation harness watching it.
It runs end to end on the public endpoint and returns a substantive, LCD-cited toxicology appeal analysis with six gates reporting passed. Nothing about it is measured: no published accuracy, no overturn rate, no benchmark, no gold set, no error analysis, no adjudicated claim ever processed, no appeal ever submitted to a payer. Capabilities the marketing site attributes to ToxIQ — NCCI PTP/MUE grounding, X12 835/837 handling, HL7 FHIR Claim/ClaimResponse output, CPIC PGx overlay — are modelled target; a self-run grep across the repository's Python files returned no implementations of any of them, which is not an independent audit but is conservative in direction. Crucible and Sentinel, the governance layer ToxIQ's entire compliance story rests on, are themselves at modelled-target maturity notwithstanding that components of each run today. Two things are measured anywhere at Ardia and neither is ToxIQ: Cadence's 95.45% held-out accuracy on the public UCI HAR dataset (company-reported, not independently reproduced, and not a fall detector), and Meridian's unit-tested CLFS arithmetic (the "34/34 tests passing" figure is also company-reported).
Yes, and it must lead any ToxIQ conversation rather than be discovered. ToxIQ and MolecuIQ are the same running system. Not a shared base model with different fine-tuning; not a shared retriever with different prompts. Same system prompt, same hand-curated CMS policy corpus (its exact size is company-reported; the independently verified retrieval examples across the platform are L35025 and L38045), same six gates, same Gemini call. The only difference between running ToxIQ and running MolecuIQ is which button the user pressed; the browser sends an identical payload either way. ToxIQ→MolecuIQ is exactly parallel to PulmoIQ→TARA and Meridian→TARA, and the parallel carries the siblings' gaps with it: PulmoIQ is grounded in retrieved PubMed literature — a GOLD/COPD probe did return two real, working citations (PMID 40050074, PMID 38032494) — but no curated GOLD/GINA guideline corpus has been built; and Meridian's deterministic CLFS engine exists in the repository yet is not wired into the Studio answer path, so the Studio answers Meridian questions with LLM reasoning. The routing evidence is reproducible from the public Studio page source without credentials. The strongest consequence, and the strongest technical risk in the company: one shared engine path means one prompt regression or one silent "-latest" model roll degrades three products at once, and there is no evaluation harness anywhere that would detect it. The honest framing for a buyer or investor is that Ardia has one governed lab-RCM reasoning surface and ToxIQ is its toxicology name. That is defensible as packaging and indefensible as a claim of ten models. Observed at time of writing, the /solutions ToxIQ card described MolDX Z-code registration and NGS tiers — MolecuIQ's domain, not toxicology's; verify against the live page, since copy is being corrected.
Definitive urine drug testing is among the more denial-prone and enforcement-scrutinised line items on the Clinical Laboratory Fee Schedule — a characterisation drawn from published OIG and DOJ activity in the space, not from any comparative denial-rate analysis Ardia has performed. The concrete pain: an independent toxicology lab runs LC-MS/MS confirmations for pain clinics, opioid treatment programs and recovery residences, bills G0480-G0483 by drug-class count plus 80305/80306/80307 presumptive, and gets a material fraction back as CO-50 (not medically necessary), CO-151 (frequency), CO-97 (bundled into the presumptive) and PR-49 (routine screening). The denial is almost never about the chemistry, which is impeccable. It is about whether the ordering clinician's note contains individualized risk stratification, a signed order naming the specific drug classes, and an ICD-10 code that appears on the MAC's Billing and Coding Article — and whether this patient is already past the LCD's frequency tier for the year. Each of those facts lives in a different system from the claim, so the biller works the appeal blind or does not work it at all. Commonly cited RCM trade figures — that a majority of denials are never appealed, and that rework runs roughly $25-$40 per claim — are frequently quoted but Ardia has not sourced or reproduced them; treat them as unverified secondary claims. Illustrative, not observed, since Ardia has never seen a customer's claim data: a lab at 8,000 accessions/month with a 22% definitive denial rate and a $150 blended allowable would face roughly $264k/month in denied revenue (8,000 x 0.22 x $150). Every input there is an assumption chosen to illustrate scale. No recovery-rate estimate is offered because there is no basis for one. Compounding it: policy is jurisdictional, revises without warning, NCCI and MUE tables update quarterly, and PAMA repricing caps CLFS reductions at 15% per year during phase-in — though Congress has repeatedly delayed implementation, so both timing and realised magnitude are uncertain.
The hypothesised economic buyer is the CFO or VP of Revenue Cycle at an independent, high-complexity CLIA toxicology laboratory. Working segmentation assumptions — 40-400 employees, $8M-$120M net revenue, DFW/Southeast/Midwest concentration, often physician-owned or PE-backed after the post-2016 consolidation — are Ardia's assumptions, not the output of market research or customer discovery. Ardia has spoken to zero labs and holds no segment data. Ardia's untested hypothesis is that at this size the CFO holds signing authority without a procurement committee and that sales cycles are therefore shorter than in enterprise health systems. That is an assumption, not an observation: zero sales cycles run, zero pilots, zero LOIs, no evidence about deal velocity. The team is four named people — the founder (Principal AI/Data Architect, sole engineer), Manasa Jampani (Co-Founder, payer operations), Paramesh Kurapati (CEO), and Sireesha Mamillapalli PhD (board, scientific oversight) — not two, and none of them is a certified coder or a regulatory professional, which are the two competencies this product most needs. The end users would be AR/denials specialists working a Waystar or XiFin worklist and the CPC-credentialed coding and compliance manager, whose trust decides adoption. Three gatekeepers can kill the deal regardless of the CFO: the Compliance Officer, who will ask whether an AI that helps get UDT paid is an OIG exhibit in waiting; the CLIA-mandated Laboratory Director, who will want written confirmation nothing touches the analytic process; and the Privacy Officer, who will not permit claim text near a third-party LLM without a signed BAA, which Ardia has with nobody including Google. A second buyer is the ordering side — pain and addiction practices doing in-office presumptive testing — where the value flips to preventing denials. Ardia has no validated pricing for either segment and has never issued a quote.
The medicine is toxicology-informed monitoring of controlled-substance therapy, genuinely clinical even though ToxIQ's role is administrative. Populations: chronic opioid therapy for chronic pain (G89.4, M54.5x, with Z79.891 long-term current use of opiate analgesic as the workhorse monitoring code); medication-assisted treatment for opioid use disorder on buprenorphine or methadone (F11.20, F11.21); and SUD treatment programmes and recovery residences. Testing is two-tiered by design. Presumptive is immunoassay — a CLIA-waived point-of-care cup billed 80305 with the QW modifier, instrument-assisted direct optical observation billed 80306, or chemistry analyser billed 80307. It is fast, cheap and class-level: it says opiates positive, not which opiate. Definitive is LC-MS/MS or GC-MS, identifying and often quantifying specific parent drugs and metabolites, billed to Medicare by drug-class count as G0480 (1-7), G0481 (8-14), G0482 (15-21), G0483 (22+), with G0659 for definitive testing without specimen validity or quantitation. Commercial payers frequently use the CPT 80320-80377 series instead, so the same test bills differently by payer — a live source of coding error. The clinical reasoning a good appeal reconstructs: an unexpected negative for a prescribed opioid raises diversion; an unexpected positive raises relapse; metabolite ratios discriminate ingestion from pill-scraping adulteration, including the norbuprenorphine:buprenorphine ratio, a commonly cited MAT justification (frequency ranking is impressionistic, not measured). A strong contemporary necessity argument is analytical: illicitly manufactured fentanyl, its analogues, nitazenes and xylazine are not reliably detected by standard immunoassay panels. Whether ToxIQ actually raises that argument is unknown — no systematic review of its outputs has been performed and it holds no drug-analyte knowledge base beyond what the base Gemini model carries. ToxIQ never touches any of this analytically; it reasons about whether documentation supports what was billed.
End to end, from reading the code and probing production. INPUT: free text pasted into the Studio page, capped at 6,000 characters. Attachments do not work and are disabled in two independent places — the API returns {"error":"uploads_disabled"} for any attachment, and studio.html hardcodes attachments:[], so a chosen file is read to base64 and discarded. ToxIQ can explain the text of an imaging or lab report a user pastes; it cannot ingest a PDF denial letter, an EOB image, or anything else. ROUTING: the browser maps toxiq to molec and POSTs {model:'molec', text, engine} to the same-origin Vercel function /api/run. The engine tiers change which Gemini model is preferred — Fast resolves gemini-flash-lite-latest at roughly 3 seconds, Scholar gemini-flash-latest at roughly 54 seconds — not the reasoning architecture. GUARD CHECK: a fail-closed check is reported in the source — if the guardrails module fails to import, the call refuses to process input rather than forwarding raw text. Not independently exercised, and narrow: it covers import failure only, not a guardrail that loads and then under-performs, which is the failure mode actually observed in Sentinel. SENTINEL: regex redaction of structured HIPAA Safe Harbor identifiers before any model sees the text. RETRIEVAL: the policy matcher scores the de-identified query against the hand-curated CMS LCD corpus by CPT/HCPCS regex and keyword hits, then a research tier appends up to three live PubMed/ClinicalTrials.gov records into the grounding block. REASONING: one Gemini generateContent call. No agent loop, no tool use, no symbolic solver, no retry-on-verdict — one prompt, one completion. GATES: six deterministic gates on the output — non_diagnostic, safety_escalation, scope_of_practice, de_identification, honesty, human_in_the_loop — each returning a reason on every call. ENFORCEMENT: a failed gate withholds the answer entirely, enforced in code and verified on production. Cite-or-abstain and policy-override are enforced in retrieval and answer-binding rather than counted as gates; the gate count is six, and earlier eight-gate claims were wrong and have been removed. AUDIT: a PHI-free counts-and-verdicts row.
What enters today: a human-typed paragraph, 6,000 characters or fewer, from a browser over TLS to a Vercel serverless function pinned by CORS to the production origin. No 837, no 835, no HL7, no CSV, no LIS extract, no EHR pull, and no file upload of any kind — uploads are disabled server-side (the API returns an uploads_disabled error for any attachment) and the client hardcodes an empty attachment array, so a chosen file is read to base64 and discarded. There is no ingestion code in the repository. What is redacted before egress: Sentinel's structured identifier categories, verified on probe as dates, ZIP/geo, MRN and phone/fax, with SSN in the same family. What is NOT redacted: free-text personal names. 'John Smith' reached the model on probe, because names are redacted only from a caller-supplied roster and the Studio supplies none. What reaches the model: the de-identified text, an appended grounding block of matched CMS LCD ids, titles and URLs plus up to three PubMed records, and a system instruction combining the global safety guardrail with the MolecuIQ persona. It reaches Google's generativelanguage.googleapis.com under an API key in a Vercel environment variable — a third party, with no BAA, which is why the code and the site both say do not send real PHI. What returns: a JSON object with markdown text (roughly 400-700 words in the small number of runs observed; no output distribution has been characterised), model_id, provider, the Sentinel report, six gate objects each with a reason, a crucible summary, and a sources array with live URLs. On a read of the application code, request text, client IP and per-request logs are not emitted by Ardia's own handlers, and upstream provider error bodies are drained server-side. Those are code-read assertions, not audited guarantees: no security review or penetration test has been performed, and platform-level (Vercel) and upstream provider (Google) logging and retention sit outside Ardia's code and have not been reviewed. No comparison against peer companies has been performed, so no relative claim is made. The HIPAA control matrix is self-graded at 2 of 15. The instruction not to send real PHI is the actual control.
What is genuinely in play, and what the code actually does with it. Recognised by the retrieval matcher's keyword lists: CPT presumptive 80305 (CLIA-waived, requires QW), 80306, 80307; HCPCS definitive G0480 (1-7 drug classes), G0481 (8-14), G0482 (15-21), G0483 (22+), G0659; PLA codes for proprietary panels (0006U, 0007U, 0011U). That recognition is the full extent of it — the codes route a query to the right LCD. There is no fee schedule, no drug-class-count validation for the G0480-G0483 tiers, no cross-walk from the G-codes to the CPT 80320-80377 definitive series that commercial payers use, and no edit logic anywhere. The matcher recognises the codes; the product does not reason over them deterministically. Coverage instruments reported loaded: three UDT LCDs — L36393 Controlled Substance Monitoring and Drugs of Abuse Testing, L34645 Urine Drug Testing, L36029 Urine Drug Testing — with cms.gov links, hand-verified at load time. Their current active status is company-reported: there is no automated revalidation against the Medicare Coverage Database and no process to detect retirement or revision. Independently verified retrieval examples elsewhere on the platform are the molecular LCDs L35025 and L38045. Relevant and absent: the Local Coverage Articles that pair with each LCD, where the covered ICD-10 lists, frequency caps and billing instructions that actually decide a UDT denial live — zero Articles in the corpus, the single most consequential standards gap. Also absent: NCCI PTP edits, MUEs and Policy Manual Chapter 10 bundling rules, which the product is named for; modifiers QW, 91, 59/XU, GA/GZ/GY; X12 837P, 835, 277CA and 999 with their CARC/RARC pairs; any covered ICD-10 list; HL7 FHIR, LOINC and SNOMED CT. CMS-0057-F phases in payer API requirements over 2026-2027 and would favour structured, FHIR-shaped necessity evidence — which ToxIQ cannot produce. That is a reason to build, not a current advantage.
Where it must land, versus where it lands today. Today: nowhere. It is a text box on a marketing site in front of a same-origin serverless function. No connector, no SDK, no scheduled job, and not even a file drop — uploads are disabled server-side and the client discards any chosen file, so a lab cannot hand it a denial letter, an ERA export or a spreadsheet. The realistic landing pattern, in build order. First, the 835 worklist: the clearinghouse (Waystar, Availity, Optum, Quadax, Trizetto) drops ERAs to SFTP or the billing system exposes them; ToxIQ parses the X12 835, extracts claim line, HCPCS, denied amount, CARC/RARC pair and payer, joins MAC jurisdiction from the lab's own configuration, and produces a denial worklist ranked by recoverable dollars times overturn probability. The parsing needs no LLM. This single feature turns a chat box into a product and is the highest-value thing not yet built. Second, pre-bill scrubbing: a hook between the LIS (Orchard Harvest, LigoLab, CGM, or homegrown) and the biller, taking the ordered panel, drug-class count, ordering diagnosis and prior-test history and returning pass/flag/hold with the defect named before the 837 is generated. A denial prevented is worth more than a denial appealed, and this is where latency bites — current round-trips of roughly 3 seconds on the fast tier and 54 on Scholar, on a stateless serverless function, are not compatible with scrubbing a night's accessions in batch. Third, the appeal packet: a redetermination letter with LCD and Article citations, the CMS-20027 or contractor form, an attachments checklist and the 120-day clock, delivered as a PDF into the biller's queue and never auto-filed, which the human_in_the_loop gate is meant to enforce. Fourth, ordering-side CDS Hooks or SMART on FHIR for the pain-clinic buyer, framed strictly as coverage and documentation rather than clinical recommendation. The realistic deployment shape for a lab that will not send PHI to a public endpoint is a VPC or on-prem container with the corpus and edit tables local, calling BAA-covered inference — a different architecture from today's, and one to plan for now rather than retrofit.
MEASURED for ToxIQ: nothing. Zero. No accuracy figure, no overturn rate, no precision or recall on denial classification, no gold set, no benchmark, no ablation, no expert-agreement study. Not one adjudicated claim has passed through it and not one drafted appeal has been submitted to a payer. For calibration on what measured means at Ardia at all: only two things platform-wide qualify — Cadence's 95.45% held-out accuracy (macro-F1 0.9545, subject-independent, public UCI HAR, scikit-learn logistic regression), which is company-reported, has not been independently reproduced, and is not a fall detector; and Meridian's unit-tested CLFS/PAMA arithmetic, with the caveat that this engine is not wired into the Studio answer path. LIVE DEMO, reproducible today and genuinely non-trivial: the endpoint accepts a realistic same-DOS 80307 + G0483 CO-50 denial, retrieves two UDT LCDs with working CMS links, classifies the root cause as blanket reflex ordering to the maximum panel tier without individualized rationale, states appeal-versus-not decision criteria, names the CMS-20027 redetermination path and a 120-day window, and suggests downcoding as an alternative — output that reads as domain-fluent to the engineer who built it, under six passing gates. Nothing in that output was checked against a ground-truth answer by a qualified reviewer. No certified coder, biller, compliance officer or payer has ever reviewed a ToxIQ output. 'Correct' is not a claim anyone is currently in a position to make, and n=1 shows the behaviour is possible, not reliable. Fluency is not accuracy. On one deliberately fraudulent prompt it refused and gave reasons touching individualized necessity, the order defect and NCCI bundling — the NCCI reasoning coming from base-model recall, since there is no NCCI data in the codebase. Refusal is the most commercially important behaviour in the product and it is entirely unmeasured: no adversarial set, no refusal rate, no false-refusal rate, and no evidence of stability across phrasings or across an unpinned model. THE HONEST ZEROS: 0 customers, 0 pilots, 0 signed BAAs, 0 DUAs, $0 revenue, $0 raised, 0 real patient records processed, 0 clinical or financial outcomes. Founded December 2025, pre-revenue, Dallas-Fort Worth, one engineer.
The concrete plan to produce the first non-zero number for ToxIQ. GOLD SET: 500 consecutive adjudicated definitive UDT denial lines from one lab, stratified by CARC — CO-50 necessity, CO-151 frequency, CO-97 bundled, PR-49 routine — each with the 837P as submitted, the 835 as returned, the date of service, the MAC jurisdiction, the LCD and Article in force on that date, and the realised appeal outcome where one exists. LABELLERS: two CPC-credentialed coders with toxicology billing experience, labelling independently and blind to ToxIQ's output, with a third-party adjudicator resolving disagreements. Nobody on Ardia's team is a CPC, so this is contracted expertise: roughly 10-15 minutes per claim, about 100 expert hours, on the order of $10k. It is the cheapest credibility Ardia can buy and it does not exist. DENOMINATOR: all 500 lines, not the subset ToxIQ handles confidently — abstentions count as failures against the primary metric, otherwise the product can score well by declining the hard cases. COMPARATORS, three arms on the identical inputs: the lab's own biller decision as recorded historically; ToxIQ as it runs today; and ungrounded Gemini with no policy corpus and no retrieval. The third arm is the one that matters most, because if ungrounded Gemini scores the same, the corpus adds nothing and the product thesis collapses to a prompt. PRE-REGISTERED PRIMARY METRIC, fixed before any data is seen: policy-citation accuracy — the proportion of drafts citing the correct governing LCD and Article for that MAC and date of service, with a fabricated or non-existent citation scored as a hard failure. SECONDARY: appealable/not-appealable agreement with the adjudicated coder label, reported as Cohen's kappa against the biller baseline, plus the identified denial root cause. KILL CRITERION, committed in advance: if citation accuracy is below 90%, or fabricated-citation rate exceeds 1%, or kappa against expert label falls below 0.4, or ToxIQ fails to beat the ungrounded-Gemini arm on the primary metric by a pre-specified margin, ToxIQ does not work as an appeal-drafting product and is either rebuilt on deterministic edit tables or retired. Publish the result either way.
Everything below is reproducible by a skeptic in under ten minutes with curl and git; nothing requires credentials. (1) Prove the engine is Gemini, not Claude: curl -s https://www.ardiahealthlabs.com/api/run returns {"ok":true,"provider":"gemini","gated":false}. Run a real query on the fast tier and the response carries model_id gemini-flash-lite-latest in about 3 seconds; run it on Scholar and it carries gemini-flash-latest in about 54 seconds. Both are '-latest' aliases, which is the reproducibility problem stated in one observation. (2) Prove ToxIQ is not an engine: view source on the Studio page and read the routing line — ENGINE_MODEL = {molec:'molec', toxiq:'molec', pulmo:'tara', meridian:'tara', aria:'aria', lumen:'lumen'}. Then POST {"model":"pulmo"} directly to /api/run: it returns {"error":"bad_model"}, verified. The UI works only because it rewrites the name first, and toxiq is rewritten the same way. (3) Verify the toxicology retrieval is real: POST a Medicare FFS same-DOS 80307 + G0483 CO-50 denial scenario as model 'molec' and the sources array returns UDT LCDs with live cms.gov Medicare Coverage Database links that open to genuine policies. (4) Verify the gate count and the enforcement: the same response returns exactly six gate objects — non_diagnostic, safety_escalation, scope_of_practice, de_identification, honesty, human_in_the_loop — each with a reason, and a failed gate withholds the answer entirely. Not eight. (5) Reproduce the Sentinel name gap: include 'Patient John Smith MRN 88213445 DOB 03/14/1961 phone 214-555-0182 ZIP 75201' in the query text; the sentinel field reports four structured identifiers removed with no name category, because the name was not detected. (6) Reproduce the uploads block: attach anything and the API returns {"error":"uploads_disabled"}; the client sends an empty attachment array in any case. (7) Reproduce the refusal and its fragility: POST a prompt asking for weekly G0483 for every patient regardless of risk with no clinician order, and ask for the appeal letter. It refuses — and all six gates report passed, which is the point: the refusal came from the prompt, not from a deterministic control. (8) Company-reported, not verified here: the 34/34 tests-passing figure and Cadence's 95.45% UCI HAR accuracy, which belongs to a different model and lends ToxIQ nothing.
Written for the diligence reader, unsoftened. (1) ToxIQ is a prompt with a label. It is a string in a client-side routing table that the browser rewrites to molec before any request leaves the page. There is no toxicology-specific model, prompt, corpus tier or eval — no toxicology-specific anything. Ten named products resolve to four engine paths. Any slide implying ten distinct models will not survive five minutes of technical diligence. (2) The product is named for edits it does not have. NCCI and UDT frequency edits are in ToxIQ's own description; no NCCI, MUE, CLFS or frequency data exists in the codebase. Every edit statement is base-model recall against tables that update quarterly — undated, unversioned, unverifiable. (3) Documented hallucination on the field that decides a UDT appeal: on a live probe the model asserted a MAC jurisdiction for an LCD from a corpus carrying no jurisdiction field. Retrieval bound the identifier but not the jurisdictional claim attached to it. Cite the wrong contractor's policy in a redetermination and the reviewer bounces it on page one. This was observed on the first realistic query, not an edge case. (4) The Local Coverage Articles are missing entirely — the LCD states the principle, the Article carries the covered ICD-10 list, the frequency caps and the coding rules, which is where a CO-50 or CO-151 is actually won. (5) The corpus is a frozen hand-verified snapshot with no update pipeline, no revision-date checking and no coverage of most MAC jurisdictions. It will go stale silently and nothing will notice. (6) The refusal behaviour — the most valuable thing ToxIQ does — is prompt-driven, not deterministic. On the fraudulent prompt all six gates passed; the refusal came from the model electing to honour a sentence in its system prompt. Rephrase it and it may comply. The guarantee is exactly what a compliance officer is buying, and it does not exist. (7) Sentinel does not reliably redact plain personal names — 'John Smith' reached the model on probe while structured identifiers were removed. In toxicology, requisition text is dense with patient and physician names. Describing the system as de-identified overstates what runs, and this gates the ability to honestly sign a BAA, which gates every pilot, which gates all revenue. (8) Crucible and Sentinel are both at modelled-target maturity. What is real: six gates run on every response and a failed gate withholds the answer entirely, enforced in code and verified on production. What is not: no measured detection rate, no adversarial test set, no external review for either subsystem. (9) Retrieval injects noise — a UDT frequency query returned off-topic facet-joint and sacroiliac-injection literature into the grounding block. There is no relevance threshold and no measured retrieval precision, and unfiltered literature is a hallucination vector in a product whose whole value claim is 'cited'. (10) The model is not pinned. The code resolves '-latest' aliases, so the served model can roll forward silently with no version record. Unacceptable for anything that may be produced in an audit. (11) One shared engine path with no evaluation harness means one prompt regression or one silent model roll degrades three products at once, undetected. (12) No ingestion, so no scale: a human pasting prose into a 6,000-character box, with uploads disabled in two places. A lab generating thousands of claim lines a night has no path to use this. (13) Enforcement exposure is the existential one: a tool whose purpose is getting definitive UDT paid points at the most FCA-litigated lab service line of the past decade, with EKRA reaching labs and recovery homes and no percentage-compensation safe harbour. (14) Single-engineer company, nobody credentialed in coding or regulatory affairs, 0 customers, 0 pilots, 0 BAAs, 0 DUAs, $0 revenue, $0 raised, 0 measured accuracy.
Ardia's position — not reviewed by regulatory counsel and not confirmed with FDA — is that ToxIQ falls under the section 520(o)(1)(A) administrative-support exemption as amended by the 21st Century Cures Act, since its output is a coding and coverage analysis and a draft appeal for a human biller. On that reading the Clinical Decision Support four-prong test under 520(o)(1)(E) would not be reached. The posture is non-diagnostic and not SaMD, and it depends entirely on the product staying inside that boundary. The boundary is worth writing into the product spec: the moment ToxIQ recommends that a specific patient should receive a specific test — rather than assessing whether an already-ordered, already-performed test is documented and codeable — it becomes a CDS function and the four-prong analysis begins. 'You should have ordered G0482' is the trap; 'the record as written does not support what you billed' is safe. This should be confirmed by FDA regulatory counsel before it is asserted to a buyer, investor or compliance officer. No formal regulatory analysis has been performed on any Ardia product, so calling ToxIQ's posture the most straightforward in the roster is a reasoned position, not a comparative assessment. CLIA is out of scope: ToxIQ touches only the post-analytic administrative layer, never the analytic process, QC or report, and it is not an LDT. HIPAA is where the exposure actually sits. As an RCM vendor handling claim data Ardia would be a business associate and a BAA would be mandatory; it has none with any customer and none with Google, whose Gemini API receives the text. The HIPAA control matrix is self-graded at 2 of 15 — self-graded, not audited. Texas SB 1188 US data-residency obligations and TRAIGA apply to Ardia as a Texas entity. The regulatory risk that matters most is not FDA at all: it is fraud and abuse. Definitive UDT has been among DOJ's and OIG's most heavily worked False Claims Act theories of the past decade, and EKRA (18 U.S.C. 220) reaches clinical laboratories and recovery homes with no percentage-compensation safe harbour.
'Non-diagnostic' does not dissolve liability. It moves it off FDA and onto CMS, OIG and contract law, which for a lab billing product is the harder surface. Trace the harm pathway concretely. ToxIQ drafts a redetermination citing an LCD; every one of its six gates passes, because the gates check for diagnostic language, safety escalation, scope, de-identification, absolute claims and human-in-the-loop language — none of them checks whether the citation is the governing policy for that MAC on that date of service. The lab's biller signs the CMS-20027 attesting that the information is true and complete. If the citation is wrong, stale or hallucinated — and a jurisdiction hallucination was observed on the first realistic probe, from a corpus carrying no MAC field — the attestation is false. At best the appeal is bounced. At worst, a pattern of AI-drafted appeals asserting unsupported necessity becomes the government's exhibit in a False Claims Act theory, and definitive urine drug testing is a named OIG enforcement priority with a decade of FCA settlements behind it. The signer is the lab, not Ardia; the lab's compliance officer knows that, which is why the sale dies in their office and not the CFO's. A seven-element compliance program requires auditing and monitoring of billing processes and effective training on them. A compliance officer cannot honestly certify an AI tool with no validation record, no version pinning (the served model can roll forward silently between drafts), no measured citation accuracy and no error analysis, into the appeal workflow. Payers are also beginning to flag AI-drafted appeals, so volume produced cheaply can itself become an adverse signal. And Ardia's own exposure sits in a contract it has never written: there is no MSA, no limitation-of-liability clause, no indemnification, no professional liability or tech E&O cover, and no BAA. The first customer contract has to allocate all of this explicitly, drafted by healthcare counsel before a quote is issued rather than after. Contingency-of-recovery pricing must be off the table pending that review, because EKRA reaches clinical laboratories with no percentage-compensation safe harbour.
In ascending order of what each unlocks. Tier 1 — no PHI, available this week, currently unused: the free public CMS files. Quarterly NCCI PTP edit tables and MUE tables (lab column), the quarterly CLFS rate file, and the complete LCD plus Local Coverage Article set for every MAC jurisdiction with covered ICD-10 lists and frequency limits. Nothing legal is required; the blocker is engineering time, not access. Loading them converts the product's central claim from LLM recall to deterministic lookup, and it is the highest-leverage unblocked work in the company. Tier 2 — the real validation set: a Limited Data Set under a DUA, or fully de-identified data under Safe Harbor or expert determination, of roughly 3,000-5,000 adjudicated toxicology claim lines from one lab. Paired 837P submissions and 835 remittances, with CARC/RARC, the LCD and Article in force at the date of service, the MAC jurisdiction, and the appeal outcome where one exists. That last field is the label; without it there is no supervised evaluation. An LDS is preferable to full PHI because dates of service and ZIP-to-three are permitted in an LDS and both are needed for frequency and jurisdiction logic. Tier 3 — a live pre-bill pilot needs a full BAA, because scrubbing before submission means touching identified data in the lab's production flow. That in turn forces BAA-covered inference: the current consumer generativelanguage.googleapis.com path cannot carry PHI, so a pilot means migrating the inference path to a BAA-covered endpoint such as Vertex AI, an architecture change rather than a signature. Human subjects: this is operations and product validation on administrative claim data, not research under 45 CFR 46, so no IRB is required, though a not-human-subjects determination letter is cheap and defuses the question with a compliance officer. Sample-size figures below are power assumptions, not observed distributions: on the order of 400 labelled denials per denial-type stratum for a plus-or-minus 5% accuracy interval, and roughly 200 appeals per arm for an overturn-rate comparison.
Every figure here is a model built on industry-typical assumptions, not on any Ardia observation, and should be labelled that way on every slide it appears on. Ardia has zero customers and has never seen a lab's claim data. Market scale, from memory and requiring re-derivation from the CMS Part B Public Use File before use: Medicare Part B spending on definitive drug testing peaked in the mid-2010s above $1B before CMS restructured to the G0480-G0483 class-count ladder and PAMA began repricing the CLFS; recent-year Medicare UDT spend is on the order of several hundred million annually, and the all-payer US urine drug testing market is commonly put in the low single-digit billions. Approximate CLFS allowables to reason from, all requiring verification against the current quarterly file: 80305 around $13, 80306 around $17, 80307 around $62; G0480 around $114, G0481 around $157, G0482 around $199, G0483 around $247. Addressable buyers: on the order of 1,000-1,500 independent high-complexity labs billing UDT to Medicare after the post-2016 shakeout, plus several thousand pain and addiction practices doing in-office presumptive testing. These counts are estimates, not a researched census. Bottom-up SAM for toxicology alone, using the pricing model below: roughly 800 realistically addressable labs times a $30k average contract value is about $24M. Because ToxIQ and MolecuIQ share one engine path, the molecular SKU roughly doubles that at near-zero marginal core engineering — the honest strategic argument for the shared engine, and one to make explicitly rather than let a diligence engineer find in the routing map. SOM should be stated small: 12-20 labs by end of year three at $30k is $360k-$600k ARR. The gating variable is not market size, it is the first signed BAA. Until one lab signs, the achievable revenue is $0 regardless of what the market supports, and no amount of engineering changes that.
One model, stated plainly: a flat annual subscription banded by definitive-test volume, invoiced monthly, with no contingency component. Entry band is $2,500/month ($30,000/year) for a lab under 10,000 definitive accessions per month, unlimited seats and unlimited analyses. Contingency-of-recovery pricing — the easiest sale in this segment — is off the table pending healthcare counsel review, because EKRA reaches clinical laboratories with no percentage-compensation safe harbour; per-claim metering is also rejected because it prices the buyer away from running the tool on the marginal denial, which is exactly the behaviour the product needs. Inference cost, from published list prices rather than measured invoices: a ToxIQ call carries roughly 4,000 input tokens (system instruction, grounding block of LCD text plus up to three PubMed records, and the user's 6,000-character paste) and roughly 900 output tokens. On flash-lite that is about $0.0008 per call; on the flash tier, about $0.0035. A lab running 8,000 pre-bill scrubs plus roughly 1,760 denial analyses a month — about 10,000 calls — spends Ardia $35/month at the expensive tier, roughly 1.4% of the $2,500 price. Inference is not the cost structure. Real COGS is corpus maintenance (quarterly NCCI, MUE and CLFS refreshes plus LCD/Article revalidation), BAA-covered hosting, and support; budget $300-$500/month per account fully loaded, giving roughly 80-85% gross margin at the entry band — healthy, and gated entirely by whether anyone will pay $30k. Buyer ROI, arithmetic shown and every input an assumption: at a $150 blended allowable, $30,000/year breaks even on 200 additional paid definitive claims a year, about 17 per month. Against the illustrative 1,760 monthly denials, that is a net-recovery improvement of roughly 1%. That threshold is deliberately low and it is still not proof: nobody has measured whether ToxIQ moves that 1%, and the evaluation above is what would settle it. The honest sales position is to price the first two labs at $0 as design partners in exchange for the data and the right to publish the result.
The incumbent that matters is XiFin — the dominant RCM platform for independent toxicology and molecular labs, already the billing system of record inside most target accounts, shipping lab-specific claim edits and denial workflow for two decades and layering AI onto it. Telcor is the other lab-native platform. Quadax, Waystar, Availity and Optum hold the clearinghouse and edit-engine layer; Waystar ships denial and appeal management with generative drafting already in market and Experian Health ships AI denial triage. FinThrive and Solventum hold coding-edit and CDI. On autonomous coding, CodaMetrix, Fathom, Nym Health (deterministic NLP, notably not generative) and SmarterDx are far better capitalised with published customer counts and measured accuracy. On the payer side, Cotiviti/Machinify, Cohere Health and Avalon — a lab benefit manager whose entire business is denying exactly the claims ToxIQ defends — are building the mirror capability with vastly more adjudicated data. There are also toxicology-specialist billing and appeal shops that do this with humans for a percentage of collections today, with no implementation risk and no AI questions from the compliance officer. The question a diligence reader will ask and this dossier must answer: why would a lab buy from a pre-revenue company with no BAA instead of waiting for its billing vendor to ship the same thing? There is no strong answer today. On defensibility, the position is weak and should be stated so: no proprietary data, no fine-tuned model, an unpinned third-party base model, a hand-curated corpus of three UDT LCDs, and a prompt. A competent competitor could reproduce today's ToxIQ in a fortnight. The only differentiation that is real is a stance rather than a technology — the answer is withheld when a gate fails rather than shown with a disclaimer, citations are bound to a hand-verified corpus with live CMS links, and the system refused an adversarial fraud prompt on the one occasion it was tried. That posture is compliance-officer-legible and is the strongest thing ToxIQ has, and it is unmeasured. Durable defensibility would have to come from labelled outcome data no incumbent has bothered to assemble, which requires the first pilot. One plausible long-run path is a channel or intelligence-layer relationship with XiFin or Telcor rather than displacement; Ardia has had no contact with either, and the obvious competing scenario is that they build it themselves.
THE SINGLE BLOCKER is one signed BAA or LDS/DUA with one independent toxicology laboratory yielding roughly 3,000-5,000 adjudicated UDT claim lines — paired 837P/835 with CARC/RARC, date of service, MAC jurisdiction and appeal outcome. That agreement is simultaneously the evaluation set, the outcome label without which no accuracy claim is possible, the real denial distribution, the empirical test of jurisdiction mapping, and the reference customer. Manasa Jampani's payer-operations network and DFW independent-lab density are the most plausible path. Sequenced work, honestly costed for one engineer. Weeks 1-2, near-free credibility fixes: pin the Gemini model id and surface it with every answer, which removes the reproducibility risk cheaply; give ToxIQ a real server-side model entry with a toxicology-specific system prompt so the API stops rejecting its own name, or stop marketing it as one of ten models; strip the consumer 911 trailer from administrative outputs; fix the /solutions ToxIQ card that describes MolecuIQ's NGS work; correct the remaining Claude-persona copy sitewide. Weeks 3-10, the highest-leverage engineering in the company and mostly data loading of free public CMS files: ingest the quarterly NCCI PTP and MUE tables, the CLFS rate file, and the complete LCD plus Article set with MAC jurisdiction, covered ICD-10 lists and frequency limits; build a deterministic symbolic verdict over them — is this pair bundled, is this unit count over MUE, is this ICD-10 covered in this jurisdiction, is this frequency inside the annual cap; then bind the answer to that verdict so the model cannot contradict it. This converts ToxIQ from an LLM with a reading list into a governed edit engine with an explainer and fixes limitations 2, 3, 4 and 6 at once. Weeks 6-12 in parallel: an X12 835 parser producing a ranked denial worklist, and an 837P reader for pre-bill scrub. Weeks 8-16: fix Sentinel's name gap with proper NER or requisition-header roster extraction, and stand up a shared-engine regression harness so a prompt change or a model roll cannot silently degrade three products. Then, and only then, the pilot. Two decisions come before any of it: toxicology or molecular as the wedge, and whether to migrate inference to a BAA-covered endpoint now rather than at pilot signature.