← Healthful AI home Download PDF
Pilot designed, conducted and reported by
Falcon CDI mark
Case study · Shadow pilot findings

Falcon CDI shadow pilot

What a point-of-care documentation tool surfaces on real hospital records in the Kingdom of Saudi Arabia, measured against an independent credentialed coder.

Pilot completed August 2026 Kingdom of Saudi Arabia ICD-10-AM 10th Ed · AR-DRG V9.0
01 · Executive summary

What was tested, and what it showed

Healthful AI ran a shadow pilot of Falcon CDI over 23 de-identified episodes supplied by a private hospital group in the Kingdom of Saudi Arabia: 10 de-identified discharge summaries (inpatient) and 13 de-identified outpatient notes. Falcon is a point-of-care clinical documentation improvement (CDI) tool. It reads the clinical note as it is written and surfaces suggestions to the treating clinician for the specific detail that coding and classification will later depend on.

To measure Falcon, an independent Australian credentialed coder first coded the same records. The coder's documentation queries, written in the formal audit query format Australian health fund audits accept (the IHACPA guideline format), define what a professional retrospective audit would raise on each record. The pilot then measured how much of that audit value Falcon surfaces on the document itself, before any auditor is involved.

Every Falcon run in this case study was captured from server logs with each chart cryptographically bound to its logged run, and all but two charts were run repeatedly, so findings are reported as frequencies rather than single outputs.

Headline result

On the inpatient set, the independent coder's review moves the DRG on 7 of 10 episodes on the coder's own post-query coding; in six of them the movement comes from a documentation query being answered; across the eight inpatient charts measured with repeated runs, Falcon raised a suggestion against 21 of the coder's 27 query items (78%), 15 of them outright and 6 partially, leaving 6 missed. On the outpatient set, the coder amended the hospital's submitted coding on 12 of 13 visits, and Falcon raised a suggestion against 26 of the 31 amendment and query items (84%), 18 outright and 8 partially, leaving 5 missed. The two cohorts are measured against different baselines and are never combined.

Beyond the coder's own list. Falcon also raised findings that fell outside the independent coder's baseline entirely. Twelve such findings were identified across the corpus, and all twelve were put to the coder for review. Four were validated as carrying clinical evidence worth querying, and two of those moved the episode to a higher complexity level: one on a neonate in neonatal intensive care, from minor to major complications, the other from intermediate to major. Section 05 sets them out. This is a separate measurement from the headline above and the two are never combined. Two of the four carry no DRG consequence at all, and they matter for a different reason: an incomplete record is relied on by the next clinician who treats the patient.

From here. Section 09 sets out the path a shadow pilot of this kind takes into a formal pilot: a technical session with hospital IT, extension of the same measurement to a larger corpus, test environment access, and conversion to a formal pilot on the hospital's own records.

10 + 13
Inpatient + outpatient episodes
9 to 11
Runs per chart
7 of 10
Inpatient episodes move DRG (coder)
100%
Of runs surfacing each DRG driver Falcon caught
02 · The pilot

What the pilot was

The pilot compared independent views of the same records.

  1. The records

    23 de-identified episodes supplied by the hospital group: 10 de-identified discharge summaries (inpatient) and 13 de-identified outpatient notes. All records were de-identified before any processing.

  2. The benchmark

    An independent Australian credentialed coder reviewed every record. On the inpatient records the coder worked blind from the record as documented, producing a baseline DRG, a set of documentation queries in the IHACPA audit format, and the post-query DRG achievable if each query is answered. On the outpatient records the coder audited the hospital's submitted coding and recorded amendments. The coder is a benchmark, not ground truth: where Falcon and the coder diverge, this case study presents the divergence as an open question for adjudication, not as an error on either side.

  3. The system under test

    Falcon CDI was run over the same records the coder received, surfacing suggestions for the clinical specificity each record lacks.

A shadow pilot is a retrospective proxy for point-of-care behaviour

This pilot ran Falcon over completed records, which is what makes a controlled comparison against an independent coder possible. In live use, the same suggestions appear while the clinician is writing the note. That distinction matters operationally: documentation requests raised after the fact are routinely set aside, while a suggestion raised during writing reaches the clinician at the moment the record can still be improved and the clinical detail is still current. The pilot measures what Falcon surfaces; live deployment moves that moment to the point of care.

The comparison logic

The coder's queries define what documentation was missing from each record. Falcon is scored on whether it surfaces a suggestion for that same detail. If the clinician acts on the suggestion, the coder codes the improved record, and the movement from baseline DRG to post-query DRG under AR-DRG Version 9.0 is the result. The evidence unit is query-level convergence: for each coder query, did Falcon raise a suggestion for it.

Two cohorts, two distinct comparisons

  1. Inpatient

    The hospital supplied no inpatient coding for these records, so the baseline is the independent coder's own first pass. The inpatient result is therefore the CDI-answerable uplift on a competent baseline, which is the improvement available even when the coding is already professionally done. It is not a claim about the hospital's inpatient coding, and none is made.

  2. Outpatient

    The hospital's own submitted coding was available for these visits, so the outpatient cohort additionally compares against that coding. The outpatient result is expressed as coding accuracy and completeness rather than DRG movement.

Classification context

All coding is to ICD-10-AM 10th Edition, grouped with AR-DRG Version 9.0, the hospital's current classification framework, shared with Australian practice. That shared classification is what makes an independent Australian credentialed coder directly portable as a benchmark for records from the Kingdom.

03 · Method integrity

How the measurement was made

The measurement standard applied in this pilot is the one we would want applied to us.

  1. Evidence is captured from server logs, with cryptographic binding

    Every Falcon run reported here is taken from server logs, and every chart is bound to its logged run by a cryptographic hash of the record text. No result in this case study rests on a screen capture alone. A run that cannot be tied to its exact source record is excluded rather than approximated.

  2. Findings are frequencies over repeated runs, not single outputs

    Generative systems vary between runs. Charts were therefore run repeatedly, nine to eleven runs each with the two exceptions disclosed below, and every finding is reported as a frequency. Throughout this document, 8/10 means the suggestion appeared in 8 of that chart's 10 runs. No catch is quoted without its frequency.

  3. Outright, partial and missed

    Each coder query is scored at one of three levels.

    ScoreApplies when
    OutrightFalcon raised the specific detail the query targets
    PartialFalcon raised the right area, usually as an indication question against a medication or an order, without asserting the coded detail itself
    MissedNo run raised it

    On the obstetric emergency, for example, the anaemia was reached through the transfusion order in 10/10 runs while the haemorrhage assertion appeared in 2/10, so that query scores partial. Partials are reported separately throughout and are never counted as catches.

The instrument was hardened during the pilot

Where re-measurement under the hardened instrument produced a different reading from an earlier pass, the harder reading is the one reported, including where it is lower.

HardeningWhat changedEffect on results
Suggestion capacity and rankingEarly runs surfaced a fixed number of suggestions in document order. The limit was raised and ranking enabled mid-pilot.Charts run under the early limit were re-run on the improved build before scoring.
Rule-base groundingFalcon's retrieval of its CDI rule base was verified and corrected mid-pilot.Charts run before the correction were re-run on the corrected build.
Evidence captureScoring moved from on-screen capture to server-log capture with hash binding of chart to run.All reported results are log-verified.
Repeat-run samplingRun-to-run variability was quantified, and repeat sampling was introduced in response.Findings are reported as frequencies with the run count stated; several early single-run readings softened on re-measurement and the softer numbers are the ones reported.

Disclosed exception

Two of the ten inpatient records exceed a storage limit in our capture pipeline, so those two charts (the older adult intensive care admission and the older adult general medicine episode with critical care) are reported from a single fully captured run each rather than a sampled frequency. Their tallies are kept separate from the repeated-run figures and marked in the results tables.

04 · Results

Results by cohort

How to read the results

  1. The spine of the inpatient result is DRG movement under AR-DRG Version 9.0.

    Stated as classification change. Where a percentage change in reimbursement appears, it is an illustrative valuation only. No figure in this case study is a revenue claim.

  2. The DRG movement column is the coder's own post-query coding.

    The ceiling that answering those queries reaches. Falcon's contribution is the coverage column, and the uplift is addressable uplift, contingent on clinician response. The pilot does not measure clinician response rates.

  3. Coverage is measured against one credentialed coder's single independent pass.

    Not against adjudicated truth. Divergences are reported as open questions. Every quoted catch carries its run frequency.

4.1 Inpatient cohort

Inpatient headline

Seven of the ten inpatient episodes move DRG on the coder's own post-query coding. Falcon's coverage of the coder's query items is set out chart by chart below; the two single-run charts (the older adult intensive care admission and the older adult general medicine episode with critical care) are tallied separately and not blended into it.

The misses are not scattered. Of the six items Falcon did not surface, five share one cause: the only evidence for them sits in narrative prose rather than against a medication, a result or an existing diagnosis line. Those five are the adhesions on the elective repeat caesarean section, ventilation coding on the infant in paediatric intensive care, the device-complication driver on the spinal surgery episode, therapy start and stop times on the neonate in neonatal intensive care, and the malnutrition item on the cardiology episode. Section 07 sets out the mechanism and section 08 sets out what we are doing about it, where it is the first item on the development path. The sixth miss is a coding-convention code rather than a documentation gap.

EpisodeClinical settingDRG, baseline → post-query (coder)Falcon vs coder queriesNotable, with run frequencies
Inpatient 1Infant, paediatric intensive care, airway surgeryE02B → E02A4 caught · 3 partial · 1 missed of 8 (10 runs)The post-operative oedema driver behind the DRG movement raised as a medication-indication question in 10/10 runs; organism specificity 8/10; ventilation coding 0/10.
Inpatient 2Elective repeat caesarean sectionO01C → O01B2 caught · 1 missed of 3 (11 runs)Iron-infusion indication 11/11; the anaemia driver via the drug hook 11/11; peritoneal adhesions (narrative only) 0/11.
Inpatient 3Adult, ENT (nasal and sinus)D63B → D65B
(principal-dx re-sequencing)
1 caught of 1 (10 runs)Antibiotic indication asked for both antibiotics in 10/10 runs, the corpus's most stable catch. The movement itself rests on the coder's principal-diagnosis selection.
Inpatient 4 †Older adult, intensive care admissionA14B → A14B
(no movement)
3 caught · 8 not surfaced of 11 (single run)Lesion primary-versus-metastasis and septic-shock specificity raised at question level; ventilation timing and nutrition items absent in the single run.
Inpatient 5Adult, cardiology (electrophysiology)F24B → F24A2 caught · 1 missed of 3 (10 runs)Heart failure off the elevated natriuretic peptide: key 8/10, question layer 10/10; obesity 3/10; the malnutrition complexity item 0/10.
Inpatient 6 †Older adult, general medicine with critical careF05A → F05A
(complexity within same DRG)
2 caught · 3 partial of 5 (single run)Transfusion-to-anaemia link caught via the order line; pericardial effusion with its treatment caught; organism and staging detail partial.
Inpatient 7Obstetric emergency, pretermO01B → O01A1 caught · 1 partial of 2 (10 runs)Anaemia due to blood loss via the transfusion hook 10/10; the haemorrhage assertion itself 2/10.
Inpatient 8Neonate, neonatal intensive careE70B → E41B1 caught · 1 partial · 1 missed of 3 (10 runs)The renal-fullness finding behind the DRG movement keyed and queried in 10/10 runs; therapy start and stop times 0/10.
Inpatient 9Older adult, general medicineB02B → B02B
(no movement)
2 caught · 1 partial of 3 (10 runs)Acute stroke versus the stale “sequelae” label: question layer 10/10, asserted as a diagnosis 5/10; stroke-mechanism stenosis 6/10.
Inpatient 10Adult, spinal surgeryI09C → I09B2 caught · 2 missed of 4 (10 runs)Constipation treated but never diagnosed 10/10; the device-complication driver behind the DRG movement 0/10, the pilot's clearest miss.

† Reported from a single fully captured run; see the disclosed exception in section 03. Single-run absences are reported as “not surfaced”, not as capability verdicts.

The size of the inpatient documentation gap

The coder's review moves the DRG on 7 of 10 inpatient episodes, six of them through a query being answered. Priced against the published CHI schedule, answering the queries would increase reimbursement for the documented care by 28% on average across the ten inpatient episodes in the sample: the sum of the price differences over the sum of the baseline prices for all ten episodes, with the three episodes whose DRG does not change counted at zero. Inpatient 3's movement is also counted at zero, because it comes from the coder's principal-diagnosis selection rather than from a query. The prices are taken at the “Medical city and tertiary type provider” tier. The hospital's own contracted tier was not available for this pilot, so that tier is a stated assumption; the pricing basis is set out below. That is the documentation gap on ten episodes, measured as the difference between the record as written and the record with the coder's queries answered.

This is the coder's own post-query ceiling and is realisable by retrospective audit. It is stated here as the target Falcon works toward at the point of care, not as a Falcon result. Falcon's contribution is the coverage column above, which is measured per query item and is not a share of this increase: individual queries carry very different classification weight, and one episode accounts for 42% of the total price difference that answering the queries produces.

The pricing basis. Two caveats attach to every percentage in this case study. The hospital's own contracted tier was not available to us, so every percentage is calculated from prices at the CHI “Medical city and tertiary type provider” tier as a stated assumption rather than from the hospital's rate; and the schedule is a draft CY 2021 price list. The tier assumption does not move the percentages: the schedule's three tiers are fixed multiples of one another for every DRG, so each movement gives the same percentage at every tier. The draft CY 2021 caveat stands, because a price list with different relative prices between DRGs would give different percentages. The classification movements underneath them are derived from the coding, not from the price, so any provider can apply its own contracted values to the same movements and every DRG finding here stands unchanged. This average excludes the further movement on the neonate in neonatal intensive care reported in section 05, which lies beyond the coder's own coding. That episode is a single path, E70B to E41B here and E41B to E41A in section 05: the two steps are sequential, not additive, and are never added together. The obstetric emergency's O01B to O01A is counted once in the 28% here; section 05 reaches the same DRG by an independent route and does not add to it.

4.2 Outpatient cohort

The outpatient comparison is against the hospital's own submitted coding for the same visits, expressed as accuracy and completeness. There is no ambulatory grouper in the Kingdom, so no DRG axis is reported for this cohort.

Outpatient headline

The independent coder amended the hospital's submitted coding on 12 of 13 visits. Falcon's coverage of the coder's 31 amendment and query items is set out below, presence-counted across 9 to 10 runs per visit.

EpisodeCoder review of submitted codingFalcon coverageNotable, with run frequencies
Outpatient 1Amended: uncoded cholesterol finding added1 caught of 1Caught but intermittent: present in 4/10 runs.
Outpatient 2Amended: symptom specificity and diabetes context1 caught of 1Pain-specificity ask 10/10; the symptom-site question is routed to adjudication.
Outpatient 3Amended: symptom and iron-deficiency coding; queried the un-indicated iron infusion1 caught · 1 partial of 2Iron-infusion indication asked in 10/10 runs, wording confirmed on the clinician-facing screen; the worked example in section 06.
Outpatient 4Amended: vitamin and mineral deficiency coding3 caught of 3Order-line indication asks for each supplement 10/10, with an identical suggestion set in every run.
Outpatient 5Amended: joint and abdominal pain specificity2 caught · 1 partial of 3Breath-test and ultrasound indication asks 10/10.
Outpatient 6Amended: gastritis with organism, deficiency and obesity coding1 caught · 1 partial · 1 missed of 3The organism causal chain 0/10 and the obesity item 0/10; the least stable chart in the corpus (2 to 23 suggestions per run).
Outpatient 7Amended: rosacea and vitamin D deficiency1 caught · 1 partial of 2Acne-medication indication ask 9/9.
Outpatient 8Amended: hypothyroidism type, iron deficiency, menstrual irregularity1 caught · 1 partial of 2Iron-deficiency ask 9/9.
Outpatient 9Amended: septum and turbinate coding1 caught · 1 partial of 2Test-axis ask 8/10; symptom axis 3/10.
Outpatient 10Amended: sinusitis site, septum and turbinate, procedure coding1 caught · 1 partial · 1 missed of 3Deviated septum 10/10; turbinate 0/10; the acute-versus-chronic question is routed to adjudication.
Outpatient 11Amended: dermatitis, corn and callus, procedure coding2 caught · 2 missed of 4Per-order indication asks 9/9; the dermatitis diagnosis line 0/10 in both instances.
Outpatient 12Amended: aftercare and device-complication coding2 caught · 1 partial · 1 missed of 4Debridement and analgesic indication asks 9/9; the aftercare item 0/9.
Outpatient 13Amended: dry-eye coding1 caught of 1Dry-eye specificity 9/9 with exact wording; see section 05 for this visit's find beyond the audit.
05 · Beyond the audit ceiling

Findings only Falcon could make

A retrospective audit, however good, can only find what is already in the record. A point-of-care tool can cause information to enter the record, because it asks the clinician while the patient is still under their care and the details are still current. An auditor or coder raises the same question at least a week later, when the episode is closed and the clinician's recall of it has faded. Any finding in this class sits above the ceiling of retrospective review.

How these findings were selected. Falcon surfaces many suggestions on each record. Those suggestions were scored into findings and compared against the independent coder's baseline. Most aligned with the coder, and that alignment is what section 04 measures. This section covers the remainder: every finding Falcon raised that fell outside that baseline entirely. Twelve findings met that description across the corpus, and all twelve were put to the coder for review. This is the complete set, not a selection from it.

The result

Of the twelve Falcon-raised suggestions put to the independent coder, four were validated as carrying clinical evidence worth querying. Two of those four moved the episode to a higher complexity level.

#RecordWhat Falcon raisedRunsClassification effect
1Neonate, neonatal intensive care (Inpatient 8)A daily topical antibiotic with no documented indication and no site of application10/10E41B to E41A, minor to major complexity.
An increase of 104% over the E41B price, if the query is answered*
2Obstetric emergency, preterm (Inpatient 7)A treated but unnamed coagulopathy: plasma and platelet transfusion with tranexamic acid, and no coagulopathy diagnosis recorded6/10O01B to O01A, intermediate to major complexity.
An increase of 46% over the O01B price, if the query is answered*
3Outpatient 13Indication asks on two retinal imaging orders, on a visit whose coded record carries none of the findings those orders imply9/9No DRG axis on the outpatient cohort
4Outpatient 2An indication ask on a mammography order10/10No DRG axis, as above

* Percentage change over the starting DRG's price in the CHI schedule, calculated at the “Medical city and tertiary type provider” tier. The hospital's own contracted tier was not available for this pilot, so the tier is a stated assumption; the schedule's tiers are fixed multiples of one another, so the percentage is the same at every tier. The schedule is a draft CY 2021 price list. Any provider can apply its own contracted values to the same classification movements, which do not change. Each movement is contingent on the query being answered, which is the nature of a documentation query. One DRG applies per episode, so the neonate's movements are stages on a single path, E70B to E41B in section 04 and E41B to E41A here; they are sequential, not additive, and are never added together. The obstetric emergency is different: the coder's queries in section 04 and the coagulopathy finding here are two independent routes to the same O01A. Its 46% is the same movement already counted once in the section 04 average, not an addition to it.

Falcon on the neonatal intensive care record: the plan line apply fucidin locally is highlighted in the record, and the question asks what skin condition or diagnosis the topical fusidic acid was applied for, offering topical fusidic acid for skin lesion, skin infection, impetigo, eczema with secondary infection, or other. The patient banner is redacted.
Exhibit 1 · The neonate in neonatal intensive care, the topical antibiotic. Falcon keyed the plan line apply fucidin locally and asked, in 10 of 10 runs, “What skin condition or diagnosis was the topical fusidic acid applied for?” No indication and no site of application are documented anywhere in the record. This is the finding the independent coder validated, and the one that carries the episode from E41B to E41A. The record context is visible in the same frame: day 21 of life, gestational age 38 weeks, admitted with RSV bronchiolitis. The patient banner and the note date have been redacted for publication; nothing else in the frame is altered.

Accuracy is not only a reimbursement question. Two of the four validated findings carry no DRG consequence at all, and they matter for a different reason. A record that does not say why a medication was given, or why an image was ordered, is an incomplete account of the patient's care. That record is relied on by the next clinician who treats them. Capturing the record accurately is as important as the reimbursement attached to it, and an incorrect or incomplete medical record can lead to harm. Every one of the four findings above is, first, a gap between what was done for the patient and what the record says was done.

What the two DRG movements represent. Neither is a realised gain. Each is a query which, answered one way, moves the episode. On the neonate the movement sits on top of the coder's own audit, so it represents what the tool adds beyond a completed expert review. On the obstetric emergency the tool reaches the same classification the coder reached, by a different route and from different evidence, which means the movement is available on a record that receives no retrospective audit at all. Most records do not receive one.

06 · In the hospital workflow

What this means in a hospital workflow

The CDI officer's view, during the admission

For a CDI program, the operational question is whether the deficiencies worth querying the treating physician about are visible while the patient is still admitted, not months later. The pilot evidences exactly that query-surfacing workflow: Falcon's suggestions repeatedly matched the independent coder's audit-format queries (the IHACPA guideline format, the Australian audit-query convention), in several cases near-verbatim. What an auditor would raise retrospectively appeared as a suggestion on the document itself.

Worked examples, from the scored runs:

  1. Outpatient 3.

    The coder's audit query: no indication documented for the iron infusion. Falcon, in 10/10 runs, asked “What diagnosis or clinical indication is documented for the administration of FERINJECT 50MG/ML VIAL?”, quoted exactly as Falcon produced it and with the wording confirmed on the clinician-facing screen and answer options that carry the same distinction the coder's amendment coded (iron deficiency with or without anaemia). A near-verbatim point-of-care pre-emption of the audit query.

  2. The infant in paediatric intensive care.

    The coder's query: oedema due to surgery, the item behind the episode's E02B → E02A movement. Falcon raised the corresponding medication-indication question in 10/10 runs, as a question on the relevant drug order rather than a proposed diagnosis.

  3. The neonate in neonatal intensive care.

    The coder's query: renal fullness suggestive of hydronephrosis, the finding behind the episode's E70B → E41B movement. Falcon keyed and queried the renal-fullness line in 10/10 runs.

One product characteristic should be read alongside the catch numbers: on several charts the load-bearing catch arrived as an indication question rather than an asserted diagnosis. Clinicians experience a question to answer, not a label to accept, and the coverage figures are reported with that form disclosed.

Falcon on outpatient record 3: the FERINJECT order is highlighted in the record, and the question asks what diagnosis or clinical indication is documented for its administration, offering iron deficiency anaemia, anaemia of chronic disease, iron deficiency without anaemia, or other. The patient banner is redacted.
Exhibit 2 · Outpatient 3, the iron infusion. The coder's audit query on this visit was that no indication was documented for the iron infusion. Falcon asked, in 10 of 10 runs, “What diagnosis or clinical indication is documented for the administration of FERINJECT 50MG/ML VIAL?” The four options offered are the same distinction the coder's amendment turned on: iron deficiency anaemia separated from iron deficiency without anaemia. This is the near-verbatim pre‑emption described in section 06. The patient banner has been redacted for publication.

The question arrives with its answers already drafted

Where Falcon raises a question it also supplies the answer set, and both are captured in the server logs. Across the corpus, every logged question carried one: 41% a multiple-choice list, the remainder a yes or no with a free-text follow-up. The lists are specific to the order they hang from, not generic. On the infant in paediatric intensive care the adrenaline nebuliser order carried the choices bronchodilator therapy, upper airway obstruction, acute bronchiolitis, postoperative airway oedema and other, and postoperative airway oedema is the item behind that episode's E02B → E02A movement. On Outpatient 3 the iron infusion carried iron deficiency anaemia, anaemia of chronic disease, iron deficiency without anaemia and other, in all ten runs, which is the same distinction the coder's amendment turned on.

Falcon on the paediatric intensive care record: the adrenaline nebuliser order is highlighted in the record, and the question asks what diagnosis or clinical indication is documented for its administration, offering bronchodilator therapy, upper airway obstruction, acute bronchiolitis, postoperative airway oedema, or other. The patient banner is redacted.
Exhibit 3 · The infant in paediatric intensive care, the adrenaline nebuliser. The coder's query on this episode concerned oedema due to surgery, the item behind its E02B to E02A movement. Falcon asked what indication was documented for the adrenaline nebuliser and offered postoperative airway oedema among the options. The option the coder's query was seeking is among those offered, while the patient is still admitted. The patient banner has been redacted for publication.

The clinician is therefore not composing a response to a query. They are selecting from a list of options that includes the detail classification depends on, at the moment the patient is in front of them. That is the practical difference between a retrospective query, which requires the treating clinician to be located weeks later and to reconstruct the episode from memory, and a prompt answered while the record is being written.

On some of the partially-surfaced items the answer set carried exactly the detail the coder queried, even where Falcon did not assert it. On Outpatient 7 the question of which skin condition was actually being treated was never asserted, while the option list offered rosacea against the retinoid order in 9/9 runs. On the older adult general medicine episode the treated comorbidities were never proposed as diagnoses, while the option lists named them against the medications treating them in 10/10 runs. A partial of that kind is not a near miss. It is a question raised with the relevant options already in front of the clinician.

Answer sets and the coverage figures. These answer sets are offers rather than assertions. They are assessed on whether they are clinically plausible for the case, not against what the record happens to document, and they are not counted in the coverage figures above.

As calibration: an independent clinically-informed reviewer, working without time pressure, predicted 16 of the coder's 27 inpatient query items outright, with 3 more partially, 19 in total. Falcon surfaced 15 outright, with 6 more partially, 21 of the same 27, at the point of care. The two do not catch the same items, the comparison is recall-only, and Falcon's misses include driver-grade items both the reviewer and the coder found. Read with those caveats, a point-of-care tool landing in the same range as an unhurried clinical read is the pilot's calibration result.

Falcon does not assign codes. In the deployed workflow Falcon performs documentation improvement in real time at the point of care, and coding and grouping are then carried out separately by the hospital's coders. Grouping is a coder-initiated step rather than an automatic one.

Earlier in the journey: admission and pre-authorisation documentation

The Kingdom's forthcoming DRG implementation requires pre-admission authorisation requests through Nafis carrying principal and secondary diagnoses, documented by the physician at admission. The capability demonstrated in this pilot is not specific to discharge summaries: Falcon applies the same specificity mechanism to any clinical document, including admission requests and progress notes. The suggestions shown here at discharge can equally be raised at admission, where the pre-authorisation record is created.

Specificity and first-pass denials

Research indicates that 15 to 25% of claims in the Kingdom are denied at least once, and documentation specificity gaps are among the drivers of first-pass denials under NPHIES. The outpatient cohort in this pilot shows the relevant mechanism directly: queries an auditor would later raise on submitted coding being surfaced up front, before the claim is formed. The pilot did not measure denial rates; what it evidences is the mechanism by which denial risk is reduced.

07 · Known limitations

Known limitations and product findings

Reported in full, because the measurement is only worth what its limits are understood to be.

  1. Run-to-run variability, quantified.

    Falcon's output varies between runs on the same record, and the size of the variation is chart dependent. On the same build and day, five charts returned their lead suggestion in every run at constant rank, while the least stable chart ranged from 2 to 23 suggestions per run on identical input with no suggestion common to all of its runs, and two charts produced unions of 44 and 47 distinct suggestions against intersections of 2 and 5. No single run is representative in either direction: on the older adult general medicine episode, one run carried a high-significance occlusion finding that none of the other nine reproduced; on the adult ENT episode, the full-specificity discharge diagnosis appeared in only 1/10 runs while other runs dropped whole suggestion classes. This is why every finding in this case study is a frequency over repeated runs rather than a single output.

  2. Detail held only in narrative prose is a systematic gap.

    Falcon's recall is strongest where a finding has a structured hook in the record, such as a medication, a laboratory value, or an existing diagnosis line. Findings documented only in narrative prose can be missed in every run: peritoneal adhesions 0/11 (the elective repeat caesarean section), ventilation coding 0/10 (the infant in paediatric intensive care), the device-complication driver 0/10 (the spinal surgery episode), therapy start and stop times 0/10 (the neonate in neonatal intensive care), the malnutrition complexity item 0/10 (the cardiology episode). This was identified and quantified during the pilot and is on the product improvement path.

  3. Administrative and data-quality items are out of scope.

    Falcon's suggestions target clinical documentation specificity. It does not flag administrative or data-quality issues in the record: episodes filed under a different specialty than the care described, and implausible recorded observations, drew no suggestion in any run (0 of 81 runs across the eight repeatedly-run inpatient charts). None are scored in this pilot.

  4. Precision was not measured by this instrument.

    The pilot measures Falcon's coverage of the coder's queries. Whether every Falcon suggestion merits a clinician's attention is a precision question that requires a clinician review pass, and it is not measured here.

  5. Clinician response is a mechanism, not a measured result.

    The reported uplift depends on clinicians acting on point-of-care suggestions. A suggestion raised during writing is more likely to be acted on than a query raised months later, but this pilot did not measure response rates, and no response rate is claimed.

  6. Scope of the product lane.

    Falcon addresses documentation specificity at the point of care. Retrospective error detection, including over-coding review, is a separate capability; a companion retrospective audit product (Horizon) exists and is outside the scope of this pilot.

08 · What we are changing

What this pilot changed in our development plan

A pilot that only confirmed what we already believed would not have been worth running. Most of the value of this exercise to us is in the misses, and this corpus has changed what we build next. The items below are set out so a reader can judge the direction of the product and not only its present state.

  1. Reading records exactly as they are written.

    These records carry a high volume of spelling variation and irregular formatting, considerably more than the material Falcon had previously been developed against. Falcon identifies the exact clinical text a suggestion refers to so it can be highlighted in place, and that link is harder to make reliably when the underlying text is inconsistently written. The results in this case study were produced on a build where a failure to make that link does not prevent the suggestion being shown, so the figures here are unaffected by it. What this pilot changed is that tolerance of imperfect source text is now a design requirement rather than an assumption. The standard we are working to is that the quality of typing in a record has no bearing on what a clinician is shown.

  2. Detail that exists only in prose.

    Section 07 records this as the systematic gap. Falcon is strongest where a finding has a concrete hook in the record, such as a medication, a result, or an existing diagnosis line, and weakest where the only evidence is a sentence of narrative. Several of the misses in this case study are of exactly that kind, and they include findings that carry classification weight. This is the largest single improvement available to the product. The independent coder's queries from this pilot give us a precise specification to build against, which is something we did not have before this exercise.

  3. Consistency between readings.

    As section 07 sets out, Falcon does not return an identical set on every reading of the same record, which is why every figure here is a frequency rather than a single output. Raising that consistency means a clinician sees the strongest findings on the first reading rather than on a later one. It is being addressed directly, and it lifts the floor of the results reported here rather than the ceiling.

  4. Records of this scale.

    Two of the inpatient records are larger than anything in our earlier testing, large enough to exceed limits in our own capture layer. Records of this size are now a defined requirement rather than an assumption, which is a change this pilot caused.

  5. Tailoring answer sets to the patient.

    The options offered against an order should reflect who the patient is. On a neonatal record, choices drawn from general dermatology are less useful than choices pitched at a newborn, and the same applies to sex-specific conditions. The date of birth is available to Falcon, but neonatal dates of birth are recorded inconsistently between hospitals and between fields within the same record, so we are making the patient's age and sex explicit inputs to answer-set generation rather than leaving them to be inferred.

None of these ask anything of the hospital. They are ours to close, and they arrive as improvements to the product.

09 · What follows

From shadow pilot to formal pilot

A shadow pilot of this kind is a measurement exercise, not a deployment. Where a provider chooses to take it further, the path is the same four steps.

  1. Technical session with hospital IT.

    A working session between the provider's IT team and the Healthful AI team to cover integration, hosting, and security questions.

  2. Further records.

    The measurement approach in section 03 applies unchanged to a larger corpus, and a larger corpus strengthens the frequencies reported here.

  3. Test environment access.

    Access to a test environment for the provider's own teams, as a subsequent step following the technical session.

  4. Conversion to a formal pilot.

    Subject to the outcome of the steps above, conversion of the shadow pilot into a formal pilot on the provider's own records. Scope, timeline and terms are set out separately.

Annex A · Evidence

Evidence behind the findings

This annex records what each headline claim rests on, and lists the captures from the clinician-facing screen reproduced in the body. It is provided so that the claims can be checked rather than taken on trust.

A.1 How the evidence was captured

  1. Server logs, not screenshots

    Every figure here is taken from Falcon's server logs. Each chart is bound to its logged run by a cryptographic hash of the record text, so a result cannot be attributed to the wrong record. A run that could not be tied to its exact source record was excluded rather than approximated.

  2. Repeated runs

    Each chart was run repeatedly, nine to eleven times, and every finding is reported as a frequency over those runs. Two inpatient records (the older adult intensive care admission and the older adult general medicine episode with critical care) exceed a storage limit in the capture layer and are reported from a single fully captured run each; that exception is disclosed in section 03 and marked in the results tables.

  3. Screen corroboration

    Alongside the logged runs, one attended run per chart was captured as a series of screenshots of the clinician-facing application, 269 images in total, each bound to its chart by the patient banner visible in the frame. These were used to confirm that what the logs record is what a clinician sees. Section A.4 reports the result of that comparison.

A.2 What each claim rests on

ClaimWhereEvidence
Falcon raised a suggestion against 21 of the coder's 27 inpatient query items04.1Eight inpatient charts, nine to eleven logged runs each, scored per query item against the coder's written queries
Falcon raised a suggestion against 26 of the 31 outpatient amendment and query items04.2Thirteen outpatient visits, repeated logged runs, scored against the coder's amendments to the hospital's submitted coding
The coder's review moves the DRG on 7 of 10 inpatient episodes, six through a query being answered04.1The coder's own baseline coding and post-query coding, each grouped under AR-DRG Version 9.0
Across the ten inpatient episodes, answering the coder's queries would increase reimbursement for the documented care by 28% on average04.1The published CHI price schedule applied to all ten inpatient episodes: the sum of the price differences over the sum of the ten baseline prices, the three episodes with no DRG change counted at zero and Inpatient 3's movement, which comes from principal-diagnosis selection rather than a query, also counted at zero, at the medical city and tertiary provider tier. The hospital's contracted tier was not available, so the tier is a stated assumption; the schedule's tiers are fixed multiples, so the percentage is the same at every tier, and the movements themselves are independent of price
Four of twelve Falcon-only findings validated by the coder05The coder's written response to each of the twelve, supplied after the coder had completed the independent coding pass
The question arrives with its answers already drafted06Answer sets recorded in the logs for every logged question; three reproduced as Exhibits 1 to 3, listed in A.3

A.3 The exhibits

Three captures from the clinician-facing application are reproduced in the body, each placed with the finding it evidences. All are from the attended runs described in A.1 and all records are de-identified. In each image the left pane is the record as the clinician sees it, with the text Falcon has keyed highlighted; the right pane lists Falcon's suggestions; and the panel at lower right is the question as it reaches the clinician, with its answer options and a single Submit action. The patient banner across the top of each capture has been redacted for publication. The note date in Exhibit 1 has also been redacted. No other part of any frame has been altered.

ExhibitRecordWhat it evidencesWhere
1Neonate, neonatal intensive care (Inpatient 8)The topical antibiotic with no documented indication: the finding the coder validated, carrying E41B to E41ASection 05
2Outpatient 3The near-verbatim pre-emption of the coder's audit query on the iron infusionSection 06
3Infant, paediatric intensive care, airway surgery (Inpatient 1)The option the coder's query was seeking is among those offeredSection 06

A.4 The screen agreed with the logs

Because the figures here are drawn from server logs, the screenshots were used to test whether the logs and the clinician's view agree. Across the attended charts, every suggestion recorded in the logs was also present on the clinician-facing screen. On 14 of the 21 charts compared the two matched exactly. On the remaining seven the screen additionally showed cards generated from terms previously approved elsewhere in the same organisation, which are not counted anywhere in this case study. There was no chart on which a logged suggestion failed to appear.

Answer options are assessed on whether they are clinically plausible for the case, not against what the record happens to document, and they are not counted in the coverage figures in section 04.

↑Top