ARDIA PRECISION HEALTHGoverned AI for healthcare revenue & precision care
360° view · Elder care

Cadence

Cadence is two things sharing one name: an offline 561-feature scikit-learn logistic regression that labels 2.56-second windows of waist-worn phone motion into six activity classes at a company-reported, independently unreproduced 95.45% subject-independent accuracy on the public UCI HAR benchmark, and a Google Gemini-backed chat persona in Ardia Studio that never loads that model — with no feature-extraction code connecting either one to a real sensor, and zero older adults ever scored.

● measured Engine · Two engines, and they never meet. THE CLASSIFIER: scikit-learn 1.6.1 Pipeline(StandardScaler → LogisticRegression(max_it
UCI HAR dataset licence CC BY 4.0 (Anguita et al., 2013) — attribution requiredFDA General Wellness: Policy for Low Risk Devices (2019) — the only clean non-device pathway currently available to CadenceFDA Clinical Decision Support Software guidance (2022) / 21st Century Cures §3060 / FD&C §520(o)(1)(E) — working interpretation, not counsel-reviewed: the signal-acquisition prong likely puts an accelerometer-based product outside this exemptionFD&C Act §201(h) medical device definition — the gate that blocks RPM/RTM billing for a non-deviceCPT RTM 98975, 98977, 98980, 98981CPT RPM 99453, 99454, 99457, 99458CPT 99091 (collection and interpretation of physiologic data)CPT CCM 99490, 99439; complex CCM 99487, 99489; PCM 99424–99427HCPCS G0438/G0439 (Annual Wellness Visit — functional ability and fall-risk element)HCPCS G0556/G0557/G0558 (Advanced Primary Care Management)ICD-10 M62.84 sarcopenia, M62.81 muscle weakness, R62.7 adult failure to thrive, R26.2 difficulty walking, R26.81 unsteadiness on feet, R29.6 repeated falls, Z91.81 history of fallingOASIS-E and MDS 3.0 Section GG mobility items (IMPACT Act standardized assessment; submitted via CMS iQIES)CDC STEADI fall-prevention algorithmHEDIS Fall Risk Management measure / CMS Medicare Advantage Star RatingsHL7 FHIR R4 Observation (category: activity), US Core, FHIR Personal Health Device IG, IEEE 11073-10441, SMART on FHIRLOINC 41950-7 (steps in 24 hours) and SNOMED CT concepts for ambulation and sedentary behaviourHL7 v2 ORU^R01 for legacy senior-living platformsHIPAA Privacy and Security Rules; 45 CFR 164.514(b)(2) Safe Harbor de-identification; NIST SP 800-66r245 CFR 46.102(e) — human-subjects determination for secondary analysis, made by an IRB of record, not the investigatorFTC Health Breach Notification Rule (non-HIPAA consumer app path)Texas SB 1188 (US data residency); TRAIGA / HB 149 (AI disclosure)False Claims Act and OIG seven-element compliance program guidance — the operative exposure for any output feeding a signed CMS assessment or claim

Where it sits in the platform

Elder care — pillar 5 of five. Ardia spans (1) precision medicine, (2) molecular and genomic diagnostics including the kidney SKU and PGx, (3) the 2027 PAMA rate cliff, (4) pulmonary and respiratory care, (5) elder care; toxicology sits on the roster alongside. Cadence sits in pillar 5 alone and should be positioned that way explicitly against the full set. It is the intended movement layer under Aria and Ardia One, with no code connecting Cadence to either today. Adjacency to pillar 4 — activity tolerance as an input to PulmoIQ — is aspirational on both sides: nothing in the repo links them, and PulmoIQ's own grounding is retrieved PubMed literature rather than a curated GOLD/GINA guideline corpus, which is not built. No relationship whatsoever to pillars 1, 2 or 3; Cadence is not a claims product, touches no specimen, and has no exposure to the PAMA cliff. Stating that plainly is worth more than manufacturing platform adjacency: a diligence reader who sees a movement classifier claiming relevance to molecular diagnostics discounts everything else on the page.

Status, stated precisely

MEASURED — COMPANY-REPORTED, NOT INDEPENDENTLY REPRODUCED. The company reports 95.45% held-out accuracy and macro-F1 0.9545, subject-independent, on public UCI HAR, from a trained scikit-learn logistic regression. These figures come from the company's own metrics.json and training script. No third party has re-downloaded the dataset, re-run the committed artifact, or confirmed the confusion matrix. The methodology described in the training script — train-only model selection, a hard disjoint-subject assertion, a single scoring pass — reads as clean on inspection, but "reads as clean" is not "reproduced." Supporting figures inherit the same caveat and are company-reported, not reviewer-confirmed: balanced accuracy 0.9533, three-state Active/Sedentary/Resting rollup 0.9963, worst-subject 0.8571, mean 0.9523, best 0.9948, 2,947 test windows across 9 held-out subjects asserted disjoint from 21 training subjects. Independent reproduction is an open item and it is cheap: the dataset is public under CC BY 4.0 and the 41 KB artifact is committed. Two separate labels are required for the rest of the product. LIVE DEMO: the Cadence persona in Ardia Studio, a Gemini-backed framing under six deterministic gates, with no published accuracy and no evaluation set — this is what a customer can actually touch. MODELLED TARGET: the deterioration signal, the rolling personal baseline, the anomaly detector, elder-cohort accuracy, any gait metric, and any link to Aria or PulmoIQ. Cadence is explicitly NOT a fall detector. The zeros are unchanged: 0 customers, 0 pilots, 0 signed BAAs or DUAs, $0 revenue, $0 raised, no real patient data, no clinical outcomes, and 0 older adults ever scored by this model.

Shared engine path — read this first.

Material and prominent: the Cadence a customer can touch is not the Cadence that has a number. The measured classifier is offline scikit-learn, served by nothing. The live surface is a Gemini persona — an LLM framing that never loads cadence_model.joblib and never touches the 95.45%. Cadence therefore has the same disease as the rest of the stack, in an unusually acute form: Ardia's ten named models collapse to four verified engine paths, ENGINE_MODEL = { molec, toxiq→molec, pulmo→tara, meridian→tara, aria→aria, lumen→lumen }. ToxIQ runs MolecuIQ's prompt; PulmoIQ and Meridian run TARA's. That map contains no "cadence" key, and posting an unmapped id straight to the API returns {"error":"bad_model"} — verified with model:"pulmo" — so the Studio must be rewriting the Cadence id client-side to one of the four paths before sending. Which path has not been confirmed and should not be asserted until it is. What is certain either way: the persona is a prompt with a label, and the trained artifact is never invoked. Meridian is the precedent for how badly this reads — it has genuinely deterministic CLFS/PAMA arithmetic in models/meridian/clfs.py, mirrored client-side, that is not wired into the Studio answer path, so a buyer meets two different Meridians. Cadence is the same pattern with a trained model instead of an arithmetic engine. The structural consequence is the strongest technical risk in the company: one prompt regression or one silent "-latest" model roll degrades three products at once, with no eval harness to detect it. Cadence's differentiation cannot be "we have a real model" while the shipped surface never calls it.

01

What it is, and who it is for

The problem

Functional decline in older adults is slow, silent, and invisible until it produces a catastrophe. The concrete pain, with figures that need explicit sources and as-of dates before they are used anywhere: roughly 1 in 4 Americans 65+ report a fall each year (~14M falls, ~3M ED visits, ~800k hospitalizations, ~300k hip fractures), and CDC/Florence et al. put nonfatal fall medical costs near $50B annually in 2015 dollars, of which Medicare paid roughly $29B. The clinically important fact is that the fall is the end of a process, not the start: weeks of reduced ambulation, longer sitting bouts, fewer stair transitions and more time lying down precede it. This drift is under-observed in routine care pathways. The adult daughter sees her mother weekly and normalizes gradual change. The home-care aide is present three hours a day and is neither trained nor paid to track trends. The physician gets the Annual Wellness Visit functional-ability and fall-risk element (G0438/G0439) from a self-reported questionnaire and a thirty-second observation, and self-report under-captures decline because the patient has already adapted. The plan sees nothing until a claim for DRG 480-482 arrives. So the pain is a measurement gap in reimbursed workflow. It is not, however, an unobserved signal: passive in-home sensing, wearable gait analytics and remote patient monitoring vendors all target exactly this drift today, and no competitive scan has been performed to establish otherwise. The defensible framing is a workflow gap in reimbursed care, never an empty field. And the target is decline detection, not fall detection — a crowded, different problem the company explicitly disclaims and should keep disclaiming.

Who buys it

There is no single buyer, and that is a real go-to-market problem: the person who feels the pain (the adult daughter) cannot pay for a population signal. Four candidate archetypes, ranked by desk analysis of incentive alignment — not by validated demand. Zero customer discovery interviews have been conducted; the company has 0 customers, 0 pilots and 0 signed BAAs or DUAs. (1) Risk-bearing Medicare Advantage plans and I-SNP/D-SNPs: economic buyer VP Clinical Programs or CMO, cheque signed by the plan CFO from medical-management or supplemental-benefit budget, end user a care manager working a queue. The hook is Star Ratings and the HEDIS Fall Risk Management measure, not falls avoidance, because Stars money dwarfs episode savings. (2) PACE organizations: fully capitated on Medicare and Medicaid, so the org owns 100% of the downside of a hip fracture — the cleanest ROI logic in US healthcare and the archetype whose incentives most plausibly align on paper. Untested. Small cheque, fast decision, one Medical Director rather than a committee. (3) Senior living and assisted living operators: buyer is VP Clinical/Resident Services or COO at a regional operator, per-bed-per-month opex, motivated by staffing efficiency and liability defense rather than clinical outcome. (4) Home health agencies under PDGM and HHVBP, chasing readmission and functional-improvement scores. Explicitly not a buyer today: the individual consumer. Consumer PERS is commonly priced in roughly the $25-50/month range depending on vendor and contract [source and date needed], and that spend buys an emergency summons. Cadence does not summon help. The one confident internal claim: caregivers are the end user and the churn risk, never the cheque-signer, and describing them as the buyer should stop.

Clinical & domain context

The domain is geriatric functional assessment and the frailty/sarcopenia complex, not fall medicine. The relevant constructs are the Fried frailty phenotype — slowness, weakness, exhaustion, low activity, weight loss, of which Cadence touches low physical activity and in principle slowness — plus sarcopenia (M62.84), adult failure to thrive (R62.7), generalized muscle weakness (M62.81), abnormalities of gait and mobility (R26.2 difficulty walking, R26.81 unsteadiness on feet), history of falling (Z91.81) and repeated falls (R29.6). The validated instruments Cadence must earn its place against are the Short Physical Performance Battery, Timed Up and Go, 4-metre gait speed, and CDC's STEADI algorithm. Gait speed is among the best-evidenced functional markers of decline in older adults, with <0.8 m/s marking elevated risk and <0.6 m/s high risk (Studenski et al., JAMA 2011). The uncomfortable implication should stay exactly this blunt: Cadence cannot produce it. It emits a categorical activity label — not velocity, not steps per minute, not stride length, not gait variability — despite being named Cadence. The workflows it could legitimately serve are the Annual Wellness Visit functional-ability element, home-health OASIS-E Section GG, and SNF MDS 3.0 Section GG, all of which currently rest on a single-timepoint clinician observation that continuous data could contextualize. The honest clinical claim available today is narrow and should never be widened without evidence: 'this person's proportion of active, sedentary and resting time changed relative to their own baseline.' That is a prompt for a human to look. It is not a risk score, not a diagnosis, and not a prediction of any clinical event.

02

How it actually works

Architecture, end to end

Separating what exists from what is asserted. WHAT EXISTS: (1) Input — UCI HAR's pre-computed X_train.txt/X_test.txt, 561 engineered features per 2.56-second window (128 readings at 50 Hz, 50% overlap) from a waist-mounted Samsung Galaxy S II accelerometer and gyroscope, with time- and frequency-domain statistics, jerk signals, magnitudes and FFT coefficients. (2) A hard assertion that train and test subject ID sets are disjoint, failing loudly otherwise. (3) Model selection by 5-fold StratifiedKFold CV over two candidates on the training split only. (4) Refit and a scoring pass on the 9 held-out subjects. (5) A fixed dictionary mapping 6 classes to 3 product states — walking/upstairs/downstairs to Active, sitting/standing to Sedentary, laying to Resting — where reported accuracy is highest (0.9963) because the only substantial confusion, sitting versus standing at 58 windows, is contained inside the Sedentary bucket. (6) Persistence to metrics.json and a 41 KB joblib. WHAT DOES NOT EXIST, and this is the whole gap: no signal acquisition, no sampling, no windowing, no feature extraction to reproduce those 561 features from a live sensor, no on-device or server inference endpoint, no baseline store, no anomaly detector, no alerting. The conclusion that follows and must be stated: because no feature-extraction code exists, the model can only consume a benchmark's pre-engineered features from a waist-mounted phone worn by adults aged 19-48 in a lab. Even a perfect reproduction of 95.45% would say nothing about accuracy on a real older adult, because the pipeline that would produce the input has not been written. Sentinel and Crucible sit in the LLM request path; the classifier never enters it.

What data flows where

Two entirely separate flows; conflating them is the main integrity risk. FLOW A, the classifier today: public CC BY 4.0 text files on disk to NumPy array to StandardScaler to logistic regression to an integer class label to metrics.json. Zero PHI, zero network egress, nothing to de-identify. Genuinely safe and genuinely inert. FLOW B, the live persona today: user free text to /api/run, then Sentinel de-identification, then Google Gemini (gemini-flash-lite-latest on Fast, gemini-flash-latest on Scholar), then Crucible's six deterministic gates — non_diagnostic, safety_escalation, scope_of_practice, de_identification, honesty, human_in_the_loop — each returning with a reason on every call, with a failed gate withholding the answer entirely, enforced in code and verified on production. Do not inflate the count: cite-or-abstain and policy-override are enforced in retrieval and answer-binding and are not gates. Both Sentinel and Crucible are modelled-target maturity, not production-hardened. The verified gap, stated without softening: Sentinel does not reliably redact plain personal names — 'John Smith' reached the model on probe — while structured identifiers (SSN, phone, MRN, dates, ZIP) do redact. The code is honest about this; the runtime gate is not, and a green de_identification result over a leaked name is a false all-clear that a plan privacy officer will write up as a control failure, not a nuance. Note also that attachments are disabled in two independent places — the API returns uploads_disabled and the Studio hardcodes an empty attachment list — so there is no file ingestion path at all. FLOW C, what a deployment requires and what does not exist: continuous 50 Hz tri-axial capture, on-device feature extraction, encrypted feature-vector upload, per-person baseline store, audited caregiver alerting. None designed.

Standards & policy it works to

None of the following are implemented in Cadence today. The module is three files with no ingestion, no inference service and no interoperability layer. This is a map of what a shipped product would have to speak, not a capability list. Billing instruments that could fund it: RTM CPT 98975, 98977, 98980/98981; RPM CPT 99453, 99454, 99457, 99458; CPT 99091; CCM 99490/99439 and complex CCM 99487/99489; PCM 99424-99427; Advanced Primary Care Management HCPCS G0556/G0557/G0558; and Annual Wellness Visit G0438/G0439, whose required functional-ability and fall-risk element is the most natural clinical anchor. The trap is real and important: RPM and RTM both require data from a device meeting the FD&C Act §201(h) definition, which Cadence explicitly is not, so those codes are visible but unreachable — the same non-device posture that keeps Cadence out of FDA jurisdiction also blocks the billing-enablement business model. Assessment and quality instruments: OASIS-E and MDS 3.0 Section GG mobility items under the IMPACT Act, CDC STEADI, HEDIS Fall Risk Management feeding MA Star Ratings. Diagnosis codes for context, never assertion: M62.84, M62.81, R62.7, R26.2, R26.81, R29.6, Z91.81. Interoperability: HL7 FHIR R4 Observation with category 'activity', US Core, the FHIR Personal Health Device IG, IEEE 11073-10441, SMART on FHIR, LOINC (41950-7 for steps per 24h; coverage for derived activity-state summaries is thin and local codes mapped to SNOMED CT would be needed), and HL7 v2 ORU^R01 for legacy senior-living platforms. Keep excluding X12 837/835, MolDX/DEX Z-codes and NCCI/UDT — those belong to MolecuIQ, ToxIQ and Meridian, and importing them signals a platform pitch rather than a product.

How it lands in a real customer

There is nothing to integrate yet. Say that first, then be concrete about the shape it must take. Data capture: Apple HealthKit and Google Health Connect are the realistic ingestion surfaces for a phone-based product, with on-device feature extraction so raw 50 Hz signal never leaves the handset — a privacy design and a bandwidth necessity at once. A wearable path (Apple Watch, Fitbit, or a commodity BLE accelerometer) requires a BLE/GATT stack that does not exist. Note that the current Studio has no file ingestion at all: attachments are disabled at the API and hardcoded empty in the client, so a CSV of sensor data cannot even be uploaded today. Clinical delivery: FHIR R4 Observation resources with category 'activity', LOINC-coded where codes exist, written via SMART on FHIR app launch or a backend service; Epic on FHIR and athenahealth Marketplace are the ambulatory routes. For the buyers who matter here the EHRs differ: PointClickCare is widely reported as the leading EHR in SNF and senior living, with MatrixCare and WellSky prominent in home health and hospice and Netsmart in behavioural and post-acute [sources needed]. Access to those channels appears to run through partner API programs whose eligibility requirements, costs and review timelines have not been investigated. Middleware such as Redox or Health Gorilla is a plausible way for a single-engineer company to avoid writing many point integrations; neither has been contacted, priced or scoped, and no integration partner conversation of any kind has taken place. Quality flows: OASIS-E and MDS 3.0 submit to CMS via iQIES, so Cadence would contribute evidence to Section GG items rather than transmitting directly. Legacy: HL7 v2 ORU^R01, and scheduled SFTP CSV drops for operators with no API — more of this market than vendors admit. Strike X12 837/835, clearinghouses and LIS integration from any Cadence deck.

03

Proof, and the honest state of it

Evidence today

Three tiers, applied to every component. MEASURED, COMPANY-REPORTED, UNREPRODUCED: the six-class classifier at 95.45% held-out accuracy and 0.9545 macro-F1, subject-independent, on public UCI HAR. Supporting figures carry the same caveat — 0.9533 balanced accuracy, 0.9963 three-state rollup, 0.8571 worst subject, 2,947 windows, 9 held-out subjects. Nobody outside the company has re-run the artifact. The evaluation protocol is the genuinely creditable asset here: train-only model selection, a disjoint-subject assertion, a published worst-subject floor. That protocol, not the accuracy, is what should be marketed. LIVE DEMO: the Cadence persona in Ardia Studio — a Gemini-backed framing under six deterministic gates, with no published accuracy, no evaluation set, and no measured honesty rate. This is the only Cadence a customer can touch. MODELLED TARGET: the deterioration signal, the rolling baseline, the anomaly detector, elder-cohort accuracy, any gait or steadiness metric, any link to Aria or PulmoIQ, and Sentinel and Crucible as hardened components. ASPIRATIONAL: clinical utility of any kind. On tests: the site advertises '34/34 tests passing.' That is company-reported and unverified. A larger internal count is claimed, but no reviewer has executed the suite, so no revised number should go on the site — do not replace one unverified number with a bigger unverified number. Either remove the claim until someone runs the suite and captures the output, or label it company-reported. tests/test_cadence.py is reported to assert subject-independence, dataset shape, and floors of 0.93 accuracy, 0.98 rollup, 0.80 worst-subject, plus artifact reproduction to 1e-3 — threshold gates rather than exact-value overfits, which is the better design, but unverified as passing and unverified as blocking any merge. THE ZEROS: 0 customers, 0 pilots, 0 BAAs/DUAs, $0 revenue, $0 raised, no real patient data, no clinical outcomes, 0 older adults scored.

How we will produce the first real number

The goal is one honest non-zero number about older adults. Three pre-registered studies, in order. STUDY 0, portability: re-derive UCI HAR's 561 features from the raw inertial signals shipped with the dataset and re-score the committed artifact. Denominator: all 2,947 held-out windows. Pre-registered metric: agreement with published accuracy within 1 percentage point. KILL CRITERION: if independently re-derived features drop held-out accuracy below 0.90, the artifact is not portable to any signal this company could capture and must be retrained from raw, not shipped. STUDY 1, elder generalization — the first number that matters. Gold set: 60-100 adults aged 70+, at least 30 with documented mobility impairment (gait speed <0.8 m/s, assistive device, or Z91.81), waist-worn sensor, minimum six hours of free-living wear each. Labels: two trained annotators working from a written codebook against synchronized video or thigh-worn activPAL, with Cohen's kappa >=0.80 required before any annotation counts and a third adjudicating disagreements — the labellers are contracted annotators, not the founder. Denominator: every 2.56-second window with a valid label and confirmed on-body sensor; non-wear detection must exist before enrollment or the denominator is corrupt. Comparators: majority-class baseline, HealthKit-derived activity where available, and the young-adult 95.45% as a ceiling. Primary pre-registered metric: subject-independent macro-F1 on the three-state Active/Sedentary/Resting rollup, success at >=0.90 with worst-subject >=0.80, registered on OSF before unblinding, one scoring pass, locked holdout. KILL CRITERION: three-state macro-F1 below 0.85, or worst-subject below 0.70, or the impaired subgroup more than 10 points below the unimpaired subgroup — because error concentrated on frail users is error concentrated on the entire reason the product exists. STUDY 2, the decline signal: 6-12 months, several hundred participants; pre-registered question is whether month-over-month change in active fraction adds incremental AUC over age, fall history and baseline gait speed. KILL: increment below 0.02, or a false-alarm rate above one alert per patient-month.

What a sceptic can check right now

What a skeptic can check, split honestly by who has actually run it. VERIFIED BY DIRECT PROBE on 2026-09-01: GET https://www.ardiahealthlabs.com/api/run returns {"ok":true,"provider":"gemini","gated":false}; the Fast tier reports model_id gemini-flash-lite-latest (~3s) and the Scholar tier gemini-flash-latest (~54s) — so the running engine is Google Gemini, not Claude, and the model id is unpinned. The Studio routing map collapses ten named models to four engine paths, and posting an unmapped id (model:"pulmo") returns {"error":"bad_model"}, which is why the client rewrites ids before sending. All six Crucible gates return with a reason on every call and a failed gate withholds the answer entirely. Attachments are refused at the API (uploads_disabled) and hardcoded empty in the client. Sentinel redacts SSN, phone, MRN, dates and ZIP, and does not redact a plain name — 'John Smith' reached the model. NOT RUN BY ANYONE OUTSIDE THE COMPANY, and this is the single cheapest credibility purchase available: clone the repo, download UCI HAR from archive.ics.uci.edu (61 MB, CC BY 4.0), joblib.load the committed cadence_model.joblib, predict on test/X_test.txt, and check accuracy, macro-F1, the confusion matrix, the per-subject floor and the empty train/test subject intersection against the published metrics.json. Then run the test suite and record the real count. Until someone does that, 95.45% is a company-reported number. WHAT FAILS ON INSPECTION and a skeptic should also check: listing models/ml-wellbeing returns three files, proving no feature-extraction or inference code exists; grepping the repo for accelerometer, gyro, FFT or windowing returns nothing beyond the training script. In short: the benchmark is plausible and cheap to verify, and the product wrapped around it does not yet exist.

Where it breaks

Read this section first. (1) The model cannot run on a sensor. No code anywhere transforms a raw accelerometer stream into the 561 features it requires — no resampling, no windowing, no FFT, no jerk derivation. The module is three files. Cadence can score UCI HAR text files and literally nothing else, so every claim about on-device inference or 'reads the accelerometer a smartphone already carries' describes unwritten software. (2) The live Cadence is a prompt with a label. A prospect who opens the Studio and types a question is talking to Gemini; the classifier is never invoked, and the persona borrows the classifier's credibility without using it. (3) The measured number does not transfer to the target user. UCI HAR volunteers are 19-48, waist-mounted, scripted, in a lab. Older adults have shorter strides, slower cadence, higher step-time variability, altered trunk sway and assistive devices — exactly the properties those features encode. Degradation on a frail 82-year-old is not a rounding error, and it will be worst on the subgroup that is the entire point. (4) The headline metric is not competitive: 95.45% sits at or below the ~96% the dataset's own 2013 authors published. (5) The worst held-out subject scored 85.71% — one healthy young person in nine misclassified about one window in seven. (6) The name promises a metric the model does not compute: a class label, not steps/min, gait speed, stride length or variability. (7) The deterioration signal — the entire product value — is an unbuilt heuristic with no threshold, no validation and no evidence that activity-class drift predicts anything. (8) Sentinel issues a false all-clear on plain names. (9) Non-wear is indistinguishable from inactivity, so the loudest alarm a decline detector would raise is a phone left on the kitchen counter. (10) One shared engine path plus an unpinned '-latest' model means a silent roll can degrade three products with no eval harness to catch it. (11) Single engineer, no CI on the artifact, no second reader.

04

Regulation, liability and data

Regulatory posture

Cadence's posture is administrative and decision-support software, non-diagnostic, not an FDA-regulated device and not SaMD. That is the company's own assessment, not reviewed by regulatory counsel and not confirmed with FDA. It is defensible today largely because the product performs no clinical function — the classifier is offline and has never scored an older adult — and the boundary narrows the moment capability is added. Today the classifier reads as a General Wellness product under FDA's 2019 guidance: labelling motion as walking, sitting, standing or laying promotes general fitness and references no disease. Marketing language of the form 'the earliest signal of functional decline' or 'change that can precede a fall' references a condition and drifts toward device territory; whether such copy is presently live on the Cadence page should be re-checked against the current site, which is actively being corrected, including removal of stale '8 gates' and Claude-as-engine claims. A working interpretation, not a regulatory determination and not counsel-reviewed: the signal-acquisition prong of the CDS carve-out (FD&C §520(o)(1)(E), as implemented in FDA's 2022 CDS guidance) appears to exclude software that acquires, processes or analyzes signals from a signal acquisition system. An accelerometer plausibly qualifies, which would put Cadence outside the CDS off-ramp the text-based models rely on and leave General Wellness as the only clean path — a path that forbids fall-risk scoring and frailty staging. This is significant enough to warrant paid counsel before any decline-detection claim ships. 'Not a fall detector' is therefore not merely honest positioning; it is the load-bearing regulatory wall. CLIA is inapplicable. HIPAA attaches on first plan or provider deployment. Texas SB 1188 US data residency and TRAIGA disclosure apply. The HIPAA control matrix is self-graded 2 of 15.

When it is wrong, who is holding the bag

'Non-diagnostic' does not dissolve liability; it moves it from FDA to CMS, OIG and contract law, which for anything touching reimbursement is the harder surface. Trace the pathway. Cadence's output is a categorical activity state. The moment that state is used to support an OASIS-E or MDS 3.0 Section GG mobility item, an Annual Wellness Visit functional-ability element, or an RTM/CCM time-based service, it enters a document a human clinician signs attesting it is true, accurate and complete. That attestation is the False Claims Act hook: Section GG scores drive PDGM and PDPM case-mix payment, so an unvalidated AI-derived mobility observation that inflates an assessment converts a software error into a false claim, with treble damages and per-claim penalties, and no 'the model said so' defence. The platform-level exposure is sharper still on the lab side, where urine drug testing is a named OIG enforcement priority and payers are beginning to flag AI-drafted appeals — a wrong or misapplied citation inside a CMS redetermination signed by a human is the same mechanism with a bigger number attached. A lab or agency compliance officer operating a seven-element program cannot straightforwardly permit an unvalidated AI tool into an attestation workflow; the validation and monitoring elements alone require evidence Ardia does not have. Second pathway: reliance harm. A caregiver reads 'stable' and does not escalate; the parent falls. Negligence and failure-to-warn attach to the vendor regardless of device status, and the human_in_the_loop gate is the intended liability transfer — but transfer only works if the human can actually check the output, and a bare class label with no confidence score, no non-wear flag and no provenance is not checkable. Third: the unpinned '-latest' alias means Ardia cannot reproduce the model that produced the harmful output, which is a discovery problem, not a technical footnote. Ardia has written no customer contract, has no indemnification or limitation-of-liability language, no tech E&O coverage, and a self-graded 2-of-15 HIPAA matrix while shipping a redactor that passes plain names — the exact facts a plaintiff or a plan's counsel would lead with.

What data it needs to be validated

The most under-exploited fact here: Cadence's elder-validation problem does not require a customer, a BAA or a DUA. It does require care about what public data can actually do. NHANES 2011-2014 Physical Activity Monitor data is public, includes adults 65 and over, and is the cheapest route toward an older-adult signal — but its suitability for this model is unverified and there are known obstacles. NHANES PAM is wrist-worn where UCI HAR is waist-mounted, so the 561 features do not transfer without re-derivation and revalidation, and NHANES carries no ground-truth activity labels, so it cannot score a six-class classifier directly. The honest statement is that a public-data path plausibly exists and has not been scoped, not that it is ready to run at zero friction. Other candidates and their real costs: UK Biobank accelerometry requires an approved application and access fee; FARSEEING is the reference real-world fall dataset in older adults and carries its own terms; SisFall includes older-adult participants with raw labelled signals; NHATS is the standard for disability trajectories in Medicare beneficiaries. Secondary analysis of existing de-identified datasets may qualify as non-human-subjects research under 45 CFR 46.102(e), but that determination is made by an IRB of record, not the investigator, and the company has no research compliance function to make or document it. Access agreements apply even where IRB review does not. What genuinely requires a BAA, IRB and informed consent is the only study proving clinical utility: a prospective cohort of adults 70+ wearing the sensor at home against a labelled reference standard — 60-100 participants minimum, 30+ with documented mobility impairment, 7-14 days of wear, ground truth from video annotation or thigh-worn activPAL, with SPPB and gait speed at baseline. A decline claim needs 6-12 months and several hundred participants.

05

The business around it

Market & economics

Illustrative sizing, built on unvalidated assumptions and offered as a structure to test, not a result. Population inputs each need an explicit source and as-of date: US adults 65+ (~62M, Census); Medicare Advantage enrollment (~34M, CMS enrollment files); PACE (~180 orgs, ~80,000 participants, National PACE Association); senior living (~30,600 communities, ~1.2M beds, industry census); Medicare-certified home health agencies (~11,400, CMS). Assumption A: ~15% of MA members carry a functional-decline or fall-risk flag worth monitoring — no clinical or actuarial source; this is a guess needing grounding in HEDIS fall-risk denominators or plan HCC data. Assumption B: $4 PMPM for a passive signal layer — no comparable contract, no benchmark, no pricing conversation, $0 revenue and 0 customers. On those assumptions: 34M x 15% = 5M x $4 x 12 = $240M/yr; 1.2M senior-living beds at $20/bed/month = $288M/yr; 80,000 PACE participants at $30/month = $29M/yr; home health perhaps $50-100M. Total roughly $500-650M/yr — arithmetic on assumptions, not a market estimate, and it must not travel into a deck as 'realistic.' The multi-billion aging-in-place TAM is mostly PERS hardware and should not be cited. Now the buyer's own ROI, run to its conclusion because the conclusion is unfavourable. A 50,000-member MA plan: ~25% fall annually = 12,500 fallers; ~2% of falls produce a hip fracture = ~250 fractures; at ~$40,000 first-year cost that is ~$10M/yr exposure. A heroic 5% reduction saves $500k. Charging $4 PMPM across all 50,000 members costs $2.4M/yr. Falls avoidance alone loses roughly 5:1 — and that is before noting the assumed effect size is zero-evidence today. The economic case must therefore be made on HEDIS Fall Risk Management and Star Ratings, where a $600M-revenue plan's 4.0-star crossing is worth tens of millions, not on avoided episodes.

Price, cost and margin

One model, stated once: $3.00 per monitored member per month, charged only on members the plan flags into the program, with a 12-month term, a floor of $5,000/month, and no per-alert or per-seat charges. Charging on monitored lives rather than total membership is the mechanic that fixes the arithmetic that sinks a whole-population PMPM — it aligns cost to the denominator the buyer actually cares about and removes the incentive to flag broadly. Cost side. The classifier is a 41 KB logistic regression: inference is microseconds and its marginal cost rounds to zero. The variable cost sits in the Gemini narrative layer. A Studio call runs roughly 3,000 input and 800 output tokens; at flash-lite list rates on the order of $0.15 per million input and $0.60 per million output tokens [list price, must be re-confirmed at contract], that is about $0.001 per call. Four caregiver-facing summaries per monitored member per month is roughly $0.004 per member per month — under 0.2% of a $3 price even if inference cost rose a hundredfold. Gross margin is therefore not constrained by inference; it is constrained by the human_in_the_loop review the gates require and by hosting, which is why the price floor must cover clinical review time, not tokens. A defensible target is 80-85% gross margin at scale, with the first ten customers materially below that because founder review time is uncosted labour. Buyer ROI, run honestly. A 50,000-member plan flags 15% = 7,500 monitored members: 7,500 x $3 x 12 = $270,000/yr. On episodes alone the case still fails — roughly 2,600 fallers, ~53 hip fractures, ~$2.1M exposure, and a hypothetical 5% reduction returns ~$105,000 against a $270,000 cost, on an effect size that is currently zero-evidence. The case that works is quality: $270,000 is 0.045% of that plan's ~$600M revenue, and the bonus and rebate-percentage step at 4.0 stars is worth tens of millions. Price against HEDIS Fall Risk Management and Star Ratings, and say plainly that the episode math does not clear.

Competition & honest differentiation

Start with the name, which nobody appears to have raised. Cadence (cadence.care) is a well-funded US remote-patient-monitoring company selling chronic care to this exact population, and Cadence Design Systems (NASDAQ: CDNS) holds a strong senior mark. Ardia's Cadence is a movement product for older adults colliding with an RPM company in the same buyer's inbox. That is a trademark and positioning problem that gets more expensive every month it is deferred. On the merits, the most dangerous competitor is Apple: Walking Steadiness ships free on every iPhone, derived from the same pocket-worn accelerometer, classifying users into OK/Low/Very Low fall risk, backed by a large validation study and already shareable through HealthKit. Apple has solved the signal-acquisition problem Cadence has not started, at zero marginal cost, on 100M+ US devices. Any pitch without a direct answer to 'why not just read Walking Steadiness from HealthKit?' will not survive a technical diligence call. Purpose-built passive monitoring: CarePredict (wrist ADL recognition with decline alerting — the closest analogue and considerably further along), SafelyYou (AI video fall detection, deep in memory care), Sensi.AI (audio agent, selling into exactly these buyers), VirtuSense (and notably FDA-registered, which Cadence is not), Vayyar Care and Cherish Health (mmWave radar, no wearable — a real advantage with elders who will not charge a device), Inspiren, Butlr, Origin AI. PERS incumbents own the consumer channel: Connect America/Lifeline, Medical Guardian, MobileHelp, Best Buy Health/Lively. And the quiet competitor is the benchmark itself: UCI HAR is a solved, undergraduate-accessible problem, with Anguita et al.'s own 2013 paper reporting roughly 96% on this exact split with an SVM and modern methods reaching 96-97%. Honest differentiation is emphatically not model quality. It is the governance wrapper — deterministic gates that withhold on failure, published worst-subject floors, a reproducible-from-public-data protocol — applied to elder care. That is a workflow and trust story, and a modest moat.

06

Where it goes next

Roadmap and the one unlock

Sequenced, with the dependency logic explicit and the blocker named. STEP 0, hours: correct the live inaccuracies rather than replacing them with better-sounding unverified ones — remove or label the '34/34 tests' claim until someone actually runs the suite and captures the output; re-check the Cadence page for condition-referencing marketing copy; add a hard guard in the answer-binding layer so the persona cannot render 're-validated before clinical use' as completed past tense; and resolve the naming collision with cadence.care before brand equity accrues. STEP 1, THE BLOCKER: build the raw-signal to 561-feature extraction pipeline and prove it by reproducing UCI HAR's own published feature vectors from the raw inertial signals that ship with the dataset — independently checkable, no permission required, roughly two weeks of one engineer's time. Everything downstream is gated on it. STEP 2: pin the model id. Replace '-latest' aliases with explicit versions and record the served model with every response; without this, no result is reproducible and no incident is investigable. STEP 3: build non-wear detection and a calibrated abstention path, so the model declines to score rather than reporting inactivity when the sensor is off-body. STEP 4: implement Sentinel name detection, and until it exists change the de_identification gate to report 'not assessed' rather than 'passed' — a false all-clear is materially worse than a disclosed gap, and it gates the ability to honestly sign a BAA, which gates every pilot, which gates all revenue. STEP 5: run the elder generalization study in evaluationDesign. STEP 6: ship a minimal reference app — HealthKit/Health Connect ingest, on-device features, encrypted feature-vector upload — so there is something to pilot. STEP 7: recruit one PACE organization or one regional senior-living operator as an unpaid design partner under a BAA. STEP 8: only then define the decline signal, pre-register its threshold and false-alarm target, and validate prospectively. STEP 9: take the §201(h) device question to counsel as a deliberate strategy.

07

Risks and open questions

Risk register

  • THE MODEL CANNOT RUN ON A SENSOR. No feature-extraction, windowing, resampling, FFT or inference code exists anywhere in the repo — the ml-wellbeing module is three files. Cadence can only score UCI HAR's pre-computed text files, so every product claim about reading a phone's accelerometer describes software that has not been written. Everything else is gated on this.
  • THE SHIPPED CADENCE IS A PROMPT WITH A LABEL. The live surface is a Google Gemini persona (gemini-flash-lite-latest on Fast) that never loads cadence_model.joblib. A prospect who 'tries Cadence' never touches the classifier the 95.45% belongs to, and the site's structure invites them to believe otherwise.
  • THE HEADLINE NUMBER IS COMPANY-REPORTED AND UNREPRODUCED. Nobody outside the company has re-downloaded the public dataset and re-run the committed artifact. The protocol reads clean, but until someone runs it, Cadence's single differentiator is an unverified claim — and verifying it costs one afternoon.
  • SENTINEL ISSUES A FALSE ALL-CLEAR ON PLAIN NAMES. Structured identifiers redact; 'John Smith' reached the model while the de_identification gate reported a pass. A missing name detector is a gap; a green gate over a leaked name is a control failure. This gates the ability to honestly sign a BAA, which gates every pilot, which gates all revenue.
  • ONE ENGINE PATH, NO EVAL HARNESS, UNPINNED MODEL. Ten named models collapse to four engine paths and the model id resolves a '-latest' alias, so one prompt regression or one silent model roll degrades multiple products at once with nothing to detect it — and no output can be reproduced after the fact for an incident or a legal exhibit.
  • THE MEASURED NUMBER DOES NOT APPLY TO THE TARGET POPULATION. UCI HAR volunteers are 19–48 in a lab; older adults differ in exactly the gait properties the 561 features encode. Expect a large, subgroup-concentrated accuracy drop, worst on the frailest users, who are the entire point.
  • 95.45% IS NOT COMPETITIVE. Anguita et al.'s own 2013 paper reported roughly 96% on this exact split with an SVM, and modern methods reach 96–97%. The number evidences an honest evaluation; it does not evidence a good model, and any ML-literate reader will say so.
  • APPLE SHIPS THE ADJACENT PRODUCT FREE. Walking Steadiness derives fall-risk classification from the same pocket accelerometer on 100M+ US iPhones, validated and already in HealthKit. There is no current answer to 'why not just read that from HealthKit?'
  • THE PRODUCT DOES NOT COMPUTE CADENCE. It emits a class label — not steps per minute, gait speed, stride length or variability — while gait speed is among the best-evidenced decline markers in geriatrics. The name promises the metric the model lacks.
  • THE REVENUE CODES ARE VISIBLE BUT UNREACHABLE. RPM and RTM both require data from an FD&C §201(h) device, which Cadence explicitly is not. The billing-enablement model is blocked by the same non-device posture that keeps it out of FDA jurisdiction.
  • THE CDS OFF-RAMP IS PROBABLY UNAVAILABLE. On a working interpretation not reviewed by counsel, analyzing a signal from a signal acquisition system fails prong 1 of §520(o)(1)(E), leaving General Wellness — which forbids exactly the fall-risk and decline-prediction claims the roadmap is built on.
  • FALLS-AVOIDANCE ROI IS NEGATIVE AT PLAUSIBLE PRICING. On the plan's own arithmetic the episode case loses several-fold, and the assumed effect size is zero-evidence. The case must be made on Star Ratings and HEDIS Fall Risk Management or not at all.
  • NON-WEAR IS INDISTINGUISHABLE FROM INACTIVITY. Without non-wear detection, a decline detector's strongest signal is an older adult who left the phone on the kitchen counter — the dominant false-positive mode, and it is unmitigated.
  • LIABILITY MOVES, IT DOES NOT DISAPPEAR. Any output feeding a Section GG item, an AWV element or a time-based service enters a document a clinician attests to, which is a False Claims Act surface. No customer contract, no indemnification language and no tech E&O coverage exist.
  • TRADEMARK AND POSITIONING COLLISION. cadence.care is a funded RPM company selling to the same buyers, and Cadence Design Systems holds a strong senior mark. This gets more expensive every month it is deferred.
  • SINGLE-ENGINEER KEY-PERSON RISK on a model with no serving infrastructure, no CI on the artifact and no second reader, in a company with 0 customers, 0 pilots, $0 revenue and $0 raised, founded December 2025.

Open questions — decisions still to make

  • Will the founder build the raw-signal to 561-feature extraction pipeline? This is THE blocker. Elder revalidation, the decline signal, any pilot, any billing code and any customer conversation that survives a technical question all sit behind it. It is roughly two weeks of work and requires no one's permission.
  • Will someone independently reproduce the 95.45% — download the public dataset, load the committed artifact, and publish the transcript? Until that happens the company's strongest asset is an unverified claim, and the verification costs an afternoon.
  • Will the model id be pinned? Serving a '-latest' alias means the engine can roll forward silently, no output is reproducible, and no incident is investigable. This is a governance decision, not a config detail.
  • Will the live Cadence persona be fixed or taken down? Today it borrows a classifier's credibility without invoking it. Options: hard-guard the answer-binding layer against elder-validation claims, or remove the persona until the classifier is actually served.
  • Will the de_identification gate report 'not assessed' rather than 'passed' when names cannot be checked? A false all-clear is materially worse than a disclosed gap, and no BAA should be signed while it stands.
  • Which single design partner: a PACE organization (cleanest capitated incentives, one decision-maker) or a regional senior-living operator (faster close, weaker clinical pull)? Picking one and committing beats pursuing both with zero discovery interviews conducted.
  • Is Cadence priced against Star Ratings and HEDIS Fall Risk Management, or against avoided fall episodes? The episode math does not clear at any defensible price; the Stars math does. This single choice determines the entire pitch.
  • Does Ardia accept the General Wellness ceiling permanently, or deliberately pursue FDA clearance? Decline-prediction claims and RTM billing both require becoming a regulated device. This should be a strategy with counsel, not a surprise.
  • Is the name 'Cadence' defensible against cadence.care and Cadence Design Systems — and if not, is it renamed now or after brand equity accrues?
  • Should the site stop placing 95.45% adjacent to elder-care product claims? The adjacency does the misleading, and fixing the layout costs nothing.
  • What is the non-wear detection and abstention design? Without it, the decline detector's loudest alarm is a phone on a kitchen counter, and the evaluation denominator is corrupt before the study starts.
  • Who writes the first customer contract, and what does its indemnification and limitation-of-liability language say? Ardia has never written one, and the plan's paper will arrive with uncapped PHI indemnity in it.

The other 360° views