Quantavera
Every record you depend on believes whatever it was told last.
A wrong allergy copied from one chart to the next looks exactly as trustworthy as one a doctor confirmed in person. A balance pulled from a stale statement looks as good as one the bank reported this morning. These systems grow larger every year and more confident every year. They do not grow more accurate.
Quantavera built the record that does. In an Evidence-Weighted Record, every fact keeps its source, its original document, and a learned weight that says how far that source has earned trust. Around two such records, one for a person's health and one for their money, Quantavera built the software that clinics, veterinary practices, families and financial institutions use every day.

The argument in four steps
Rounding at the door
A database decides, the moment a value arrives, whether it is true or false. Everything it knew about where that value came from is discarded at that instant and cannot be recovered later at any price.
The cliff
A record that accepts a small share of corrupted values grows more confident and less correct at the same time. Nothing in the record signals it. The more successful the record, the further its stated confidence sits from the truth.
Keep the evidence
Hold every value with its source, its date, its document and a weight learned from what actually happened. Corrections are new events pointing at what they correct. Nothing is overwritten.
Surround it with software
A record on its own sells to nobody and learns from nothing. The applications are what people buy, and they are also the supply of independent judgments the record needs to learn whom to trust.
A competitor can copy an AI scribe in a quarter. A record that knows which of its feeds broke eighteen months ago requires having kept the evidence, which is a decision made years earlier.
The suite
Two lifetime records sit at the center. Eight applications read from them and write to them, and each one is a product a company could be built around on its own.










THE ARGUMENT
More data, less truth
Every record system in health care and finance makes the same decision at the door: it rules each arriving value true or false, stores the survivors as exact, and throws away what it knew about where the value came from. A record built that way grows more confident as it grows and no more accurate.
The promise kept, and the promise broken
Share of stated 95 percent confidence intervals that actually contained the truth, as one biased source hides in a growing record.
The decision made at the door, before anyone knows the question
A value arrives. A laboratory result, a blood pressure, a bank line, a figure read off a photographed statement. The receiving system has three things it can do with it. It can accept the value, after which every report, every alert and every analysis treats the number as exactly true. It can reject the value, after which every analysis behaves as though the observation never happened. Or, most often, the value has no field to land in at all, so it goes into a free-text note or a scanned attachment where nothing can read it.
In all three cases the system rounds a continuous quantity to zero or one. How far a datum deserves to be believed is a number between those two ends, and it depends on which laboratory ran the test, whether staff confirmed the patient had fasted, whether the device had drifted, whether the amount came off a feed or out of a phone camera. The record has that information at the moment the value arrives. It discards it, replaces it with a verdict, and keeps only the verdict. Everything downstream inherits the verdict and has no way to question it.
Other fields settled this question decades ago and went the other way. A digital radio receiver can decide each bit as a hard zero or one before decoding, or pass the decoder a confidence for each bit. The soft decision is worth about 2 dB at low signal-to-noise ratios, which is the difference between a link that works and one that does not. Record systems in medicine and money still make the hard decision, at the one point in the process where the most context is available and the least is kept.
The failure that gets worse as the system succeeds
No real record is clean. Every one of them accepts some values that are wrong: a glucose drawn after breakfast and labeled fasting, a wristband reading taken while the band sat on a nightstand, a transfer between a person's own accounts tagged as income by an aggregator, a receipt photo read as $189.90 for an $18.99 grocery charge. Those errors do not cancel. A mislabeled fasting glucose runs high, not randomly high and low, so the mistakes point the same direction and leave a fixed offset in every answer the record gives.
The arithmetic is short. If a fraction of accepted values is wrong, and a wrong value is off by some average amount, the product of the two is a bias that carries no record count in it. A million records carry the same bias as a hundred. Meanwhile the stated margin of error shrinks with every new entry, because that is what margins of error do. The two curves cross. Past the crossing, the record reports narrow, confident ranges around a point in the wrong place, and nothing inside the record signals the problem. The failure is silent, and it arrives faster the better the record does at collecting data.
The size at which a record's stated 95 percent range is more likely wrong than right scales with the inverse square of the bias. Halving the bias quadruples the size a record can reach before its conclusions stop holding. Excluding weak sources lowers the bias and moves the crossing. Fixed weights, set once by a committee, lower it and move the crossing. Neither removes it, because neither drives the bias to zero. Only weights that keep learning from outcomes do that.
What the measurement shows
Three conventional methods and the evidence-weighted record were run on the same data, estimating a quantity that one source reports with a bias. Each method can be built and run by anyone. None was told which source was biased. Measured over 240 runs at each size using the shipped model, the share of stated 95 percent ranges that actually contained the truth fell to zero for every conventional method as the record grew, while the evidence-weighted record held 99 percent, and at ten thousand records every stated range held.
The pattern matters more than any single figure. The conventional curves start high at twenty-five records, where the margin of error is still wide enough to cover the bias by accident, and collapse as the record fills. A buyer evaluating one of those systems at a pilot size sees honest numbers. The same system at production size reports the same confidence and has stopped being right.
Four ways a record goes wrong, and what each one costs
Four failure modes account for nearly everything that goes wrong inside a real record. Each was simulated separately, against a conventional record and against the shipped evidence-weighted model, with no method told what it was facing.
| What is going wrong | A conventional record | This record | Difference |
|---|---|---|---|
| A source carries a hidden bias | 0% of its stated ranges held; error 2.34 | 99% held; error 0.0245 | 95x smaller error, matching the proven optimum |
| A source quietly breaks and keeps reporting | 51.5% of contested facts settle wrongly | 18.4% | 64% fewer wrong, which is 99.7% of all that was available |
| Two sources are secretly one source | 23.4% wrong | 16.1% | 31% fewer wrong, below what any weighting can reach |
| Ordinary independent mistakes | 1.63% wrong | 1.61% | Nothing, and nothing was available |
The fourth row is the one to read carefully, because it is the strongest evidence that the first three are honest. In that regime the errors are scattered rather than systematic, and a vote among independent sources is decided by who agrees with whom. Weighting every source by its true reliability, which no buildable system can do because true reliability is a property of the world, also gives 1.63 percent. The conventional record is already doing as well as anything could. This record matches it instead of degrading it. A record that claimed a general accuracy gain from learned weights in every regime would be overselling them, and the fourth row is where that claim would have had to break.
A record that accepts a small share of bad values grows more confident and less correct at the same time, and nothing inside it reports the problem.
What this costs in a clinic and in a ledger
The abstract failure has concrete forms, and anyone who has run a practice has watched them happen. A patient tells an intake clerk she reacted badly to a drug years ago. The clerk enters it as an allergy. It was a side effect of a different drug entirely, and it is now in the chart as a hard allergy. Two referrals later, three systems carry it, each having received it from a system that received it from somewhere else, and every one of them displays it as a confirmed fact with no mark of where it started. A physician who wants to use the right drug has no way to find the chain back to the clerk, and the safest action in front of them is to prescribe something worse.
In money the same failure wears different clothes. A balance goes stale: an account feed stops updating, the ledger keeps displaying the last figure it received, and nothing on the screen distinguishes a number from this morning from a number from eleven weeks ago. A loan application is prepared from it. An aggregator tags a $6,000 deposit as income when the money came out of the family's own brokerage account, and the figure lands in a debt-to-income calculation and a tax draft. A receipt photo reads $189.90 for an $18.99 charge, and the household budget now carries a category that is wrong by a factor of ten.
In each case the record had the information needed to catch the problem at the moment the value arrived: which source produced it, how that kind of source has performed before, and whether the rest of the record agreed. In each case the record decided true or false, kept the verdict, and discarded the evidence.
THE ARGUMENT
The Evidence-Weighted Record
Every fact is stored exactly once, as a permanent entry carrying what was observed, about whom, when, from which source, and how far that source has earned trust. Nothing is overwritten, and every screen, report and research set is assembled from those entries.
What a fact carries
A value in either record is never a bare number in a column. It arrives as an event and keeps what decides its meaning:
- The source, and its class. Which laboratory, clinic, device, bank, app or person produced it, and what kind of source that is: a certified laboratory, a clinic measurement, a consumer wearable, an institution feed, a figure read by software off a document, a person's own statement.
- The method. Measured, calculated, reported, or read from a document by software.
- The times. When it was observed and when it was entered, kept apart, because the gap between them is often the whole story.
- The original. A link to the document or data stream it came from, down to the spot on the page, so any number can be checked against its source years later.
- The extraction confidence. For anything read from a document or a photograph, how sure the reader was of the characters.
- The weight. How far values from this source can be believed at this kind of fact. The event keeps the weight its source carried when the value was written, and a reader also sees that source's weight now.
Weights live apart from the facts. The weight in force is not stored at all: it is computed at read time from the evidence counts, in one place both records use, so a change in policy updates every dependent analysis and neither record can drift from the other. When a weight changes, no stored fact is edited and no closed year's books are rewritten, and every analysis records the weight version it used so any result can be reproduced.
The never-overwrite rule
A fact enters once and stays. A correction is a new event that points at the one it corrects, a retraction is an event, and a merge of two records entered in error is an event, so an unmerge separates them exactly. Views show the current value; the history shows every version, who changed it, and why. Replaying the events up to a date reproduces the record exactly as it stood then, which answers the question conventional systems handle worst: what did the record say on March 3, and why did it change. A clinician defending a decision, an auditor reconstructing a filing and a lender re-examining an underwriting file all need that, and in a record that overwrites, none of them can have it.
The archive and the working record are the same object. Each event carries a cryptographic fingerprint of itself and of the one before it, and the newest is anchored outside the database every hour, so even a consistent rewrite where every internal link agrees fails against the anchors. The event tables refuse updates, deletes and truncation at the database level. Lawful erasure removes a subject's values and recomputes the chain so that it still verifies and every proof issued earlier still holds.
A reported laboratory value is never replaced or hidden by an estimate the record computed from it.
The record never overwrites a laboratory with its own opinion
The record holds a position on how far a given value can be trusted and does not act on it by changing the value. A clinician reads what the laboratory posted, as the laboratory posted it, and reads beside it what the evidence says about that source: the weight it carries at this kind of fact, how that weight was earned, and what the probability of the underlying condition looks like once the whole record is counted.
A record holds a hemoglobin A1c of 5.9 percent and earlier fasting glucose readings near 105. A new result arrives reading 132, labeled fasting, against a diabetic cutpoint of 126. With staff confirmation of the fasting draw, the value's reliability given the whole record is 0.92 and the probability the true fasting glucose is at or above the cutpoint is 0.34. Labeled fasting and unconfirmed, 0.66 and 0.25. With the patient having told the guide app she ate breakfast, 0.10 and 0.05. In all three cases the chart still shows 132, and a binary record stores it as a diabetic-range fasting value with no field in which to say otherwise.
The same rule runs through the financial record. A $61,000 deposit matched to a closed statement and an invoice reads at 0.91 reliability with a 0.67 probability of being recurring income; reported by one aggregator and unmatched, 0.42 and 0.19; read off a photographed statement with no match, 0.17 and 0.06. The deposit is $61,000 in every case, and what changes is what a loan file is entitled to conclude from it.
When sources disagree, the record settles the fact and keeps the alternatives: the value most likely to be right, its confidence, every value it rejected with that value's own probability, which sources were counted as a single witness, and the confidence the answer would have carried had those sources been treated as independent. That last figure shows what the independence assumption cost. No conventional system has a place to put it.
A registry, so a new kind of data needs no database change
In a conventional system, a new test, a new wearable metric, a genome or a new kind of account means new tables or columns, a schema migration with downtime, and changes to every program that reads them. The usual response is to push the new data into free-text notes or scanned attachments, where nobody can analyze it.
Here a new kind of data enters through a registry entry and the database structure does not change. Every event has the same outer shape; the registry defines what a valid payload for each type looks like. An entry states what the type measures, its standard code, its unit and conversions, the range of plausible values so a typo or a misread photo is held for review instead of landing, the method when that changes the meaning, its starting weight by source class, and its version, so a definition can be refined later without changing what older events meant. Adding a type takes a definition and a person's approval, and intake can accept the new data that day. The health record ships with 46 approved types.
The registry also prevents the opposite failure, a database that fills with data nobody can interpret. No value lands unless its type is defined, and anything intake cannot map is held for a reviewer rather than forced into the wrong field: a downloaded bank file with the letter O typed where a zero belonged went to the holding bay rather than into the books.
Events are the only thing written. Everything a person sees is a view computed from them and kept current as they land: the patient's timeline, the clinician's chart as FHIR resources, a research store in the OMOP format, a balance sheet. A new app or report is a new view rather than a change to what is stored.
Why one design serves health and money
Holarc holds a person's health over a lifetime. Finarc holds a person's money, and the money of every household, business, trust and estate they own, with walls between entities so figures tie together and still separate for audit. They run on one engine and one trust model, so what a weight means in a loan file is what it means in a chart, and they are never joined into one store, because a single store holding a person's complete health and financial life would be the most valuable target a thief could assemble.
The mathematics does not care what the values measure. A health record that treats a home blood pressure cuff, a clinic reading and a typed entry as equally true has the same failure as a ledger that treats a bank feed, a photographed receipt and a remembered figure as equally true. Both domains carry the same five structural problems: values rounded to true or false at entry, new data types requiring a rebuild, history overwritten, provenance lost, and one table structure serving a clinician, a patient, a researcher, an auditor and a lender equally badly.
The domains differ in one way that favors building both. Independent evidence, the kind that did not come from the system's own answer, arrives in finance every month: a statement closes, a transcript arrives from a tax authority, an instrument clears. In health an outcome can take years. The financial record is the fastest place to earn calibrated weights, and because the trust model is shared, what is learned there is not learned twice.
THE ARGUMENT
How the record learns whom to trust
A weight is a track record: how often a source has turned out to be right about a particular kind of fact, faded toward the present, counted by how independent each judgment was, and held under a ceiling when the only thing agreeing with a source is the record itself.
Reliability is learned per source and per kind of fact
A hospital that sends reliable laboratory values may send unreliable medication lists. A bank feed that reports amounts to the cent may get dates wrong. One score per source averages those together and is wrong about both. The unit of learning is therefore the pair of a source and a kind of fact, and a source excellent at amounts and poor at dates carries two weights rather than one average of them.
Dividing a source that finely would normally be fatal, because most pairs have almost no data. The estimate for a thin pair is pulled toward that source's overall rate across every kind of fact, which is pulled toward the population's rate, which is pulled toward the registry's starting figure. Splitting a source into kinds of fact therefore costs nothing where there is no data to split. No prior is allowed to become a certainty, and a weight never reaches zero or one.
The population level is built and switched off by default. With it on, a source's weight can move because some other source was judged, and the record can no longer say why that weight changed from that source's own evidence. Explainability was worth more than the small statistical gain. Learning across records belongs in the evidence engine, behind the consent gate, where a new household or practice can inherit weights learned from far more outcomes than it holds.
Evidence fades
A source that was reliable two years ago and has drifted since should not keep coasting on old credit. Counts are discounted toward the present each period. At the default rate a comparison from six months ago counts about 45 percent as much as one from this week, and a pair remembers roughly 33 weeks of evidence. Without that, a source with ten thousand good observations takes years to be marked bad, which is the failure the model exists to catch. On a feed that silently broke while its registry entry still called it good, a record that counts every judgment forever still reads it at 0.570 three years later, where a fading memory has it at 0.436 and falling.
A judgment is counted by how independent it was
A judgment that moves a weight can come from three very different places, and after the fact nobody can tell them apart unless the record wrote down which it was. It writes it down, in three stored channels never merged into one count.
Independent ground truth
A device reading against a reference instrument, a correction from a different source, a statement that closes, a transcript that arrives, an instrument that clears. The real figure became known from somewhere other than the record's own answer.
A person confirming
A clinician accepting a value may have had that value's trust weight on the screen in front of them, so the acceptance is not independent of the number it confirms.
Agreement with the record
A source agreeing with the figure the record settled on. A source that agrees with the majority is credited even when the majority is wrong.
Because the channels are stored apart, a deployment that changes what a channel is worth recomputes every weight from the judgments rather than replaying its history. A discount alone does not solve the problem: discounting self-confirming evidence by any fixed fraction changes how fast a weight climbs and not where it ends. The fix is a ceiling, which holds a weight whose evidence is mostly self-confirming to what the independent channels support, and it is one-sided, so a source agreeing only 100 times in 400 still falls to 0.27.
| Policy on self-confirming agreement | Learned weight after 400 judgments on a source worth 0.70 |
|---|---|
| Counted as independent truth | 0.952 |
| Discounted to a quarter, no ceiling | 0.934 |
| Recorded as agreement, with a ceiling | 0.773 |
Correlated sources are declared and counted once
If three clinics all pull the same figure from one regional exchange, they are not three witnesses. If two aggregator feeds draw from one bank's interface, or a feed and a statement come from the same core banking system, they are one. A product of per-source reliabilities treats them as separate, which is how a pair agreeing on the same wrong figure reads as 0.98 confident when 0.62 is warranted.
A deployment declares which sources share an upstream, and the group votes once at the weight of its most reliable member. The declaration is an asserted fact about the institutions rather than an inference from the data, because agreement alone cannot distinguish a shared upstream from two sources that are independently right. Statistical detection of copying works after roughly 48 to 220 records; declaring it is correct from the first.
Three things a conventional record has no way to attempt
Telling a biased source from a merely noisy one
A source can be bad in two ways that call for opposite responses. A noisy but honest source should carry less weight and still be used, because its errors scatter and averaging helps. A precise but biased source should be removed or corrected, because averaging it in moves the answer in one direction however many readings arrive. An agreement rate cannot tell them apart: a reading far from the truth fails the comparison either way. The sign of the errors can, and that is what the record tests, learning each source's precision from the residuals against an independent reference and taking their mean as the bias.
The test identifies the biased source and nothing else in 45 percent of runs at 100 records, 99 percent at 250, and at least 99.6 percent at every size from 500 on. Below 250 records the evidence does not yet separate a bias of 7 from noise of 9, and the test correctly stays its hand. The record's stated intervals stay honest the whole time, because the precision weights carry the protection before the removal does.
Refusing a correction that one subtraction would not fix
A source reading a constant amount high can be corrected and keeps all of its precision afterward. A source reading a fixed proportion high, or reporting in the wrong unit, cannot: one subtraction would leave it wrong at both ends of its range and right in the middle, and the fault would stop being visible. The record tests which case it faces, by the slope of the residuals as well as their mean, and refuses to correct when a constant does not explain the error. A source reading ten percent high shows a bias of 37 standard errors with a slope of 13; a source reading a constant 3 high shows 16.7 with a slope of minus 0.9, and only the second is corrected.
Counting two feeds from one upstream as one witness
A conventional record reports 0.98 confidence where 0.62 is warranted, and has no field in which to record the difference. The measurement is blunt about which tool does the work. On a correlated pair outvoting a good source, a conventional record settles 23.4 percent of contested facts wrongly; weighting every source by its true reliability, which nothing buildable can do, leaves the error exactly where it started at 23.4 percent; declaring the pair one witness takes it to 16.1 percent. Learning better weights accomplishes nothing there and the declaration accomplishes everything. Counting a family once does not recover a fact only that family witnessed: when the pair are the only witnesses and they agree, the figure stands, and what changes is that it stops reading as near-certain.
Where it buys nothing, stated plainly
Against ordinary scattered mistakes, learned weights buy almost nothing: 1.61 percent of contested facts settled wrongly against 1.63 percent for a conventional record, and weighting every source by its true reliability also gives 1.63 percent. When errors are independent and unbiased, the outcome of a vote among sources is decided by who agrees with whom, and no weighting scheme can improve on that. A deployment whose sources are all steady should expect little from the weights, and should want them anyway for the day one of those sources stops being steady.
Three further limits hold regardless. A corrupted value with no contextual signal and no conflict with the rest of the record cannot be found by any method. Weights learned without outcome data are guesses. And a weight moves only by evidence: no one sets a source's weight by hand to favor it, a manual change is a logged decision with a reason, and affiliated devices are weighed by the same rule as everyone else's.
THE ARGUMENT
What was measured
The central claim is a theorem. The gap between what a theorem makes possible and what a shipped program does was closed by measurement, against methods anyone can build, with every figure produced by the model that ships.
Median error at ten thousand records
Every bar is a method that can be built and run, shown as a multiple of the evidence-weighted result.
The theorem
For every question, every way of scoring the answer, and every prior belief, the best decisions available from an evidence-weighted record are at least as good as the best available from any binary record built from the same data, and strictly better whenever the discarded information would have changed the decision.
The proof takes three steps. The binary record is a function of the weighted one: apply the entry rule to each item and drop the rejects. Any decision rule reading the binary record therefore has a twin that reads the weighted record, computes the binary record itself, and returns the same action on every data set, at exactly the same risk. The best rule on the weighted record minimizes risk over a set containing every one of those twins, and a minimum over a larger set cannot be larger.
That is the simplest case of Blackwell's comparison of experiments from 1951: a coarsened observation is never more informative than the observation it came from, for any decision problem. In information theory it is the data processing inequality; in the philosophy of evidence, Good's principle of total evidence from 1967. There is no case in which a conventional record does better, and everything else is about how much better. The theorem applies with full force to data a binary schema never admits at all, such as a family history left in free text or cash income mentioned in conversation, where reliability was set to zero before any analysis could judge it.
A theorem says what a record makes possible. Whether the program realizes it is a separate question, and the measurements answer it for the one that ships.
Four methods, all buildable, at ten thousand records
Every comparison is against something a real system can run. Nothing was told which source was biased, and each run at a given size starts a fresh learner on a fresh history of that many records, with nothing carried over.
| At 10,000 records | Stated ranges that held | Median error | Further off than ours |
|---|---|---|---|
| Accept every value, the conventional record | 0% | 2.34 | 95x |
| Fixed weights, set once and never revised | 0% | 1.87 | 76x |
| Learned weights used only to reduce a source's influence | 0% | 1.19 | 48x |
| Holarc and Finarc | 100% | 0.0245 | the benchmark |
The third row is the method a competent team would reach for: learn each source's reliability, then use what you learned to turn the bad source's influence down. It is a real improvement over accepting everything, and it still ends 48 times further off, because turning a biased source down leaves the bias in the answer in proportion to how far down it was turned. What closes the gap is finding the source that is systematically wrong, removing it, and weighting what remains by precision. The influence-only method's stated ranges never held at any size either, so a method can be substantially more accurate than the conventional one and still report confidence that is not earned. At twenty-five records, where the margin of error is wide enough to cover a bias by accident, the three conventional methods held 65, 78 and 88 percent of their stated ranges; at a hundred, 5, 15 and 54 percent; from a thousand on, zero.
The one number above ours, and why nothing can reach it
Weight each surviving source by the inverse of its true variance and the error reaches 0.0245. Aitken proved in 1935 that this is the best any linear weighting can do, in two lines of Cauchy-Schwarz: the variance of a weighted average is minimized exactly when the weights are proportional to the inverse variances, and every include-or-exclude rule is a special case of weights set to zero and one. No binary rule built from the same readings can have lower variance.
The record reaches 0.0245. Nothing buildable can be told the true variance, because it is a property of the world rather than of any record; the bound appears as a yardstick for the weighting step, and standing on it means that step has stopped being the limiting factor. Getting there required learning the quantity the bound is stated in. Reliability and precision are different properties of a source, and a bound stated in one cannot be certified by measuring the other, so the record learns each source's precision directly, from the residuals between its readings and an independent reference. Being approximately right is cheap: the Kantorovich inequality of 1948 bounds the cost of imperfect weights, and weights off by up to a factor of about 1.4 either way cost at most 12.5 percent more variance than perfect ones, against 67 percent at any size for a binary rule that excludes weak sources carrying 40 percent of the information.
Where learning is decisive, and where the structure does the work
On a feed that was reliable and silently failed, the conventional record settles 51.5 percent of contested facts wrongly; outcome history takes that to 18.4 percent, where the best achievable by any weighting is 18.3 percent, so the record captures 99.7 percent of what was there to win. On a correlated pair outvoting a good source, the conventional record is 23.4 percent wrong, weighting every source by its true reliability leaves it at exactly 23.4 percent, and declaring the pair one witness takes it to 16.1 percent, so learning accomplishes nothing there and the declared structure accomplishes all of it. On ordinary independent mistakes, the conventional record is 1.63 percent wrong, the shipped model 1.61 percent, and weighting by true reliability 1.63 percent: nothing was available and nothing was taken.
The convergence, and the floor under it
Weights learned from outcome-linked items converge to the true values at the rate of one over the square root of the number of items that still count, which the fading memory holds to the recent past. As they converge the residual bias approaches zero, and the simulation shows the error in a learned error rate falling by roughly a factor of 3.2 for each tenfold increase in records. Underneath that sits a guarantee that bounds the downside rather than promising an upside. The family of weightings the record searches contains every binary rule, including the conventional one, and candidate weights are selected by a bounded proper scoring rule on separate validation data, which bounds how far they can fall behind the best binary rule by a margin that shrinks as validation data grow. Weights start at the conventional rule and move only as validated outcomes support the move.
What is proven and what is not
The mathematics and the behavior of the shipped model
Nine results, each derived from published mathematics: Blackwell's comparison of experiments, Good's principle of total evidence, Aitken's weighted least squares, the Kantorovich inequality, Huber and Hampel's influence analysis, Efron and Morris on shrinkage, and standard results on the consistency of learned estimates. What is new is the application to a lifetime record, the separation of judgments by their independence from the record's own output, and the measurement of what each part is worth, produced by the program that ships under a protocol stated completely enough to reproduce.
Calibration against real outcomes
No real patient record and no real financial account has entered either record. Everything runs on invented data with every outside network on a simulator. The simulations are scenarios: they check the theorems numerically, and in each one the model that learned the weights matched the process that generated the data, which a real deployment will not do exactly. A misspecified model narrows the advantage. The first calibration against real outcomes, published with its sample sizes, needs a deployment, and it is the study that converts the mathematics from proven-in-simulation to proven-in-deployment.
Three further limits are published rather than buried. A corrupted value with no contextual signal and no conflict with the rest of the record cannot be found by any method; weighting recovers nothing that was never observed. Weights learned without outcome data are guesses. And security has a boundary: an attacker who controls the running application can read what it reads for as long as the control lasts, and de-identified data is not anonymous.
The mathematics went out for an independent review
The mathematics was sent for independent review by a reader outside the company, instructed to attack it, and the review's findings are closed. A claim that has survived a hostile read is worth more than one that has not been read, and the reviewer's work is available to a prospective investor's own statistician.
THE ARGUMENT
One repository, licensed partitions
Every application in the suite reads from and writes to the same record, and none keeps a copy of its own. That removes a class of problem from the system, and it creates a second business: another company can license a partition of the same record and gain the storage, the integrity chain, the trust model and the security of the whole system without building any of it.
One record, and nothing to synchronize
Each record is the single source and repository of data for everything built on it. The patient app, the clinic system, the imaging platform and the evidence engine hold no records of their own: each writes the facts it produces into the record as events and reads back the views it needs. A laboratory result entered at a clinic and a symptom a patient reports to her guide land in the same log, under the same registry, with the same provenance fields.
There is no synchronization between applications because there is nothing to synchronize. Anyone who has run a practice on conventional software knows what the alternative costs. The practice management system has its own patient table, the portal has another, the imaging system a third, the billing clearinghouse a fourth, and an address change made at the front desk reaches two of them that afternoon, one overnight, and one never. The integration work in health care is largely the work of keeping copies of the same facts in agreement, and it is permanent, because the copies keep diverging.
Three rules keep it honest. Every finished clinical fact lives once, with images living once in the imaging platform and the record holding the report and a pointer. Every financial fact lives once, on the chain of the entity it belongs to. And applications never reach a store directly: they go through services that check role and consent on every call.
What a licensed partition is
The same repository is open to other companies under license. A licensee operates a partition inside the record: its own data, with its own encryption keys, its own trust settings, its own additions to the type registry and its own views. The walls are the ones the records already enforce between custody domains and between financial entities. Inside them the licensee gets the whole system:
- The append-only event log, with provenance on every value.
- The hash chain and its hourly outside anchors, so a rewrite inside the database is detectable even when every internal link agrees.
- The data type registry, so the licensee's own kinds of data are defined rather than forced into someone else's fields.
- The learned trust weights, with the basis channels behind each one.
- Lawful erasure that leaves every proof of the record's integrity valid, and the de-identified research path through the consent gate.
- Per-tenant, per-purpose encryption, so a stolen copy is unreadable.
Two rules govern the partitions, and they are the reason a cautious counterparty can sign. Data never crosses a wall on its own: if a person in a licensee's partition also holds a personal record, the two are linked only through the identity vault, only when the person authorizes it, and the link is a logged event. And a licensee's sources earn their weights from its own outcomes; nothing learned in one partition moves another's weights without consent on both sides.
Why that is a business
A company that produces health or financial data about people has a problem that has nothing to do with its product. It needs somewhere for the data to live that will satisfy a hospital's security review, a regulator's audit, a buyer's diligence and a plaintiff's lawyer, for as long as the data must be kept, which in health is decades. Building that is a program rather than a feature, none of it differentiates the company, and all of it has to be right.
A device or diagnostic company
Connects under commitments that protect both sides: identify the patient at the point of care, send every result in, keep its method, and be scored by the same rules as everyone else. Keeping the method is what makes it acceptable: algorithms and reconstruction methods stay on the company's own servers, and the record stores the result of using a method, never the method. Every treated patient becomes a long-term outcome record, with evidence scored by rules the company does not control, which is worth more to a regulator than evidence it graded itself.
A lender, insurer or clinic chain
A lender makes decisions from figures whose provenance it cannot see and cannot later prove it saw. A partition holds each figure with its source, the weight that source carried at the time, and the document it came from, so a file assembled in March can be reconstructed in five years exactly as it stood. A clinic chain buys for the same reason from the other direction: charts spread across acquired practices, with no way to answer what a chart said on a date.
The pricing assumption in the model is $250,000 a year for an average licensee, an assumption rather than a quoted price, and deliberately the least load-bearing line in the business at 1.4 percent of modeled 2031 revenue. The strategic value is larger, because every licensee is a source of independent outcomes, and those are the input the trust model cannot manufacture for itself.
Why it cannot be assembled from parts
Each component above exists somewhere as a product: an append-only ledger, a tamper-evident log, a key management service, a consent manager, a FHIR server, a de-identification tool. A competent team can buy all of them. What it cannot buy is a record in which they are the same object.
The integrity chain has to run over the same entries the applications read, or the archive drifts from the working record and the proof covers a copy rather than the data in use. The weights have to be computed from the same provenance fields the applications wrote, in one place, or two implementations disagree and neither can say which is right. Erasure has to be designed into the chain arithmetic from the start, because a chain built without it cannot later remove a subject's values and keep its earlier proofs valid. And the walls belong in the keys and the access rules rather than in application code, or a wall is a setting someone can change.
One engine owns what the two records share, down to the canonical encoding that gives every value one byte sequence however the database stores it, and a conformance suite covers the failures a record cannot detect from the inside, among them a consistent rewrite of the last entry that leaves every link and the stored head in agreement. Those decisions had to be made before the first application was written; a team assembling parts makes them after, in whatever order the parts allow.
The copies are what break. A record with no copies has nothing to break.
Why the suite belongs together
A record on its own sells to nobody, for a reason specific to this design rather than a general argument about bundling. The trust model learns from independent judgments: a judgment about whether a source was right that did not come from the record's own answer. Those judgments are produced by applications, in the course of work people were doing anyway. A clinic system records what a treatment did. A patient app records how someone felt three days and fourteen days afterward. An imaging platform records what a reader found and what later proved true. A financial record records whether a statement reconciled, whether a transcript matched, whether an instrument cleared.
A record surrounded by applications has a continuous supply of those judgments. A record sold on its own has none, and its weights sit near their starting figures indefinitely, which makes it a well-built conventional record with extra fields. A source that silently breaks takes a record with no outcome history to 51.5 percent of contested facts settled wrongly; outcome history takes it to 18.4 percent.
The patient app is the clearest case. Most systems treat the patient as the last person to ask, and a patient reporting what actually happened after a treatment is the scarcest input the trust model needs.
The second reason is narrower and holds regardless of the first. One sign-in, one encryption system, one gateway to outside AI, one record engine and one set of tamper-evident logs, built once rather than six times. A company that built six products on six foundations would carry six security reviews, six key management designs and six audit stories, and would be slower at every one of them.
The economics follow the same loop. A clinic running the clinic system produces outcomes. Outcomes calibrate the weights. A record whose weights are calibrated is worth more to a research sponsor and more to a lender, which raises what each will pay, which funds more clinics. That is a claim about where evidence comes from rather than a platform story, and it is checkable: it predicts that accuracy improves with the number of independent judgments reaching the record, which the simulations measure.
THE SUITE
How the parts fit
Two lifetime records, eight applications, and one foundation under all of them. Each application is a product in its own right with its own buyer and its own price. Together they are the supply of evidence that makes the records accurate.
| Product | What it is | Who uses it |
|---|---|---|
| Holarc | The lifetime health record. Every fact carries its source, its date, the original document and how far that source has earned trust. It reads laboratory results, hospital records, PDFs, faxes, phone photographs and imaging. | Everything else writes to it and reads from it. Patients never see it; their guide knows it. |
| My Talisman | The patient's own app, with a guide to talk to or type to. It answers medical questions with referenced evidence, carries the whole family, and connects to any doctor's portal. It informs; it never diagnoses or prescribes. | Patients, parents, caregivers |
| Arca | The clinic system you talk to: records and practice management run by a voice guide with people approving the work, plus the business side of the practice. | Doctors, nurses, front desk, billers, practice owners |
| Arca Vet | The veterinary edition: the client as the account with their animals as patients, veterinary terminology and formularies, the clinic pharmacy, and cash billing with full receivables. | Veterinary practices and hospitals |
| Keelson | The imaging platform: archive, viewer, radiology workflow, reporting and imaging AI. | Imaging centers and radiology groups |
| Finarc | The lifetime financial record and an adviser that does not sleep. Every account, asset, policy and obligation in one place, with walls between entities so each can be audited on its own. | Families, small business owners, and the banks, lenders and insurers that serve them |
| My Helm | The family's app for the financial record: budgets, taxes, insurance, estate planning and a standing adviser. | Households and small business owners |
| Pet Talisman | The veterinary edition of the patient app: a pet's lifetime record, preventive care and triage, carried by the family that raises it. | Families with animals |
| NERD | Numerical Evidence and Research Discipline. Scores every claim, runs research inside the walls, and returns a verdict a person can act on. | Researchers, clinics, and the guide inside every product |
| The foundation | One sign-in, one encryption system, one gateway to outside AI, one record engine, and tamper-evident logs under every product. | Everyone, invisibly |
Why they belong together, for a reason specific to this design
Every application produces outcomes, and outcomes are what the record learns from. A clinic system records what a treatment did. A patient app records how someone felt afterward. An imaging platform records what a reader found and what later proved true. A veterinary practice records what a family saw at home a week later. A financial record learns every month whether a statement reconciled.
Those are the independent judgments the trust model needs and cannot generate for itself. A record surrounded by applications has a supply of them. A record sold on its own does not, and its weights sit near their starting figures indefinitely. The second reason is narrower and holds regardless: one sign-in, one encryption system and one record engine built once rather than ten times.
One guide, every product
One guide service runs the patient's guide, the pet parent's guide, the clinic's guide and the household's financial guide, with different tools, permissions and voices in each. Patients never see the names Holarc, NERD or Keelson. Voice is never required: every spoken interaction can be typed instead, and the typing window never covers what the guide is talking about.
The line that does not move
Across the suite, AI drafts, retrieves, checks and informs, and licensed people diagnose, prescribe and sign. My Talisman never diagnoses, never prescribes beyond over-the-counter products, and never judges a photograph of a skin spot. NERD advises and people decide. A reported laboratory value is never replaced or hidden by an estimate the record computed from it: the clinician reads what the laboratory posted, and reads beside it what the evidence says about that source.

HOLARC
Holarc, the Living Health Record
Holarc keeps one record for each person, from before birth onward. Every health fact enters once, as a permanent event carrying what was observed, about whom, when, from which source and how far that source has earned trust. Nothing is overwritten and nothing is thrown away.

What a lifetime record has to survive
A person's health history is produced by institutions that never meet. A primary care office, three specialists, two hospitals, a laboratory, an imaging center, a pharmacy chain and the person themselves each hold a piece of it, in a system built to run that institution's day, keeping its own copy of the person. An address changed at the front desk reaches two of them that afternoon, one overnight, and one never. Connecting health care is largely the work of keeping copies of the same facts in agreement, and it never ends, because the copies keep diverging.
The second failure is quieter and does more damage. When a value lands in a conventional record, the system decides at that moment whether it is true and stores the winner. A clinic cuff, a wristband and a number the patient typed all become the same kind of row. The record cannot say which source it believed or what it rejected, so when a device turns out to have been drifting since January, nothing can say what it contaminated.
Holarc is built so that neither failure is possible. There is one record per person and no copies to reconcile. Every value keeps its source, its method, the time it was observed and the time it was entered, a link to the original down to the spot on the page, the confidence it was read with, and the weight its source carried at this kind of fact when the value was written. A reader also sees that source's weight now.
Events only added, views computed
Every piece of information enters as an event: this value, about this person, of this registered type, from this source, at this time, with this provenance and this weight. Events are only ever added. A correction is a new event pointing at the one it corrects, a retraction is an event, and a merge made in error is an event, so an unmerge separates the two records exactly and every view rebuilds correctly.
Everything anyone sees is a view assembled from those events: the patient's timeline, the clinician's chart as FHIR resources, a de-identified research set, a family graph, a time series for wearable readings. A view can be dropped and rebuilt from the log without touching the facts underneath, and a new application is a new view rather than a change to what is stored. Replaying the events up to a date reproduces exactly what the record said then, which answers the question conventional systems handle worst: what did the chart say on March 3, and why did it change. A clinician defending a decision, an auditor reconstructing a disclosure and a researcher barred from using later information all need that answer.
The event tables refuse updates, deletes and truncation at the database level, so an application fault or a careless administrator cannot quietly rewrite history. Working state stays out: schedules, drafts and claims in progress live in each clinic's operational store, and Holarc receives only finished facts, a signed note, a resulted laboratory test, a filled prescription.
What makes it different from a health record database
Three things, each of which a conventional record has no field in which to attempt.
Every value keeps its evidence
Provenance is part of the value rather than an audit log beside it. The source, the method, the extraction confidence and the weight travel with the number into every view and out through the standard interfaces, so an outside application reading Holarc over FHIR receives the evidence with the result.
Trust is learned per source and per kind of fact
A source good at one kind of value and poor at another carries two weights rather than one average of them. Weights are earned from outcomes, held as counts that fade toward the present, and computed at read time rather than stored, so a change in method updates every dependent analysis without editing a stored fact.
Independence is recorded, not assumed
Each judgment about whether a source was right is stored with how independent it was, and sources that share an upstream feed are declared and counted as one witness. Three clinics pulling the same figure from one regional exchange are not three witnesses, and a conventional record has nowhere to say so.
The last of those keeps the first two honest. A clinician who accepts a value may have had that value's weight on the screen in front of them, so the acceptance counts for a quarter, and a weight whose evidence is mostly self-confirming is held to what the independent channels support.
What the measurements show
Four regimes were run against the shipped model, with nothing told which source was bad, against a conventional record built from the same data.
| What is going wrong | A conventional record | Holarc | Difference |
|---|---|---|---|
| A source carries a hidden bias | 0% of stated 95% ranges held; error 2.34 | 99% held; error 0.0245 | 95 times smaller error, at the proven optimum |
| A feed quietly breaks and keeps reporting | 51.5% of contested facts settle wrongly | 18.4% | 64% fewer wrong, which is 99.7% of all that was available |
| Two sources are secretly one source | 23.4% wrong | 16.1% | 31% fewer wrong, below what any weighting alone can reach |
| Ordinary independent mistakes | 1.63% wrong | 1.61% | Nothing, and nothing was available |
The fourth row is the strongest evidence that the first three are honest. Against scattered independent error, weighting every source by its true reliability also gives 1.63 percent. The conventional record is already doing as well as anything could, and Holarc matches it rather than degrading it. A record promising a general accuracy improvement from learned weights would be overselling them.
Who it serves, and through which door
Patients never touch Holarc directly and never see its name. Each door carries a different consent and a different view.
| User | How they reach Holarc | What they see |
|---|---|---|
| Patients and caregivers | My Talisman | Their whole record, every consent they signed, a log of every read, and dependents they carry under a recorded legal basis |
| Clinics | Arca, or any EMR that sends records in | Their own chart and the outside records the patient approved them to see |
| Researchers | NERD, through the consent gate | De-identified data from people who consented, and the knowledge library |
| Connected devices | Intake, with the person identified at the point of care | The history the device needs for the patient in front of it, and its own results |
| Outside app companies | The FHIR API with SMART sign-in | Only the data their own app created |
Why everything else in the suite writes to it
The trust model learns from independent judgments, meaning a judgment about whether a source was right that did not come from the record's own answer. Those judgments are produced by applications, in the course of work people were doing anyway: a clinic system records what a treatment did, a patient app records how someone felt three days and fourteen days afterward, an imaging platform records what a reader found and what later proved true.
A record surrounded by applications has a continuous supply of those judgments. A record sold on its own has none, and its weights sit near their starting figures indefinitely, which leaves it a well-built conventional record with extra fields. That is a claim about where evidence comes from rather than a story about bundling, and it is checkable: it predicts that accuracy improves with the number of independent judgments reaching the record, which is what the 51.5 percent against 18.4 percent result measures. The narrower reason holds regardless: one sign-in, one encryption system and one record engine, built once rather than six times.
The information a conventional record rounds away at the door cannot be recovered later at any price.

HOLARC
What the record holds
Every kind of health information about a person, across a whole life, from any source, with nothing discarded. Beside the record sits a store of graded medical knowledge that holds no patient data at all.
A whole life, starting before birth
A Holarc record can begin before the person does. During pregnancy the record is the mother's, and at birth the child's record is created already linked to hers, so the child starts life with family history in place and the link never depends on the temporary name a hospital gives a newborn. From then on the record grows with the person, and nothing is removed except by lawful erasure.
A record that spans a life is a different object from a record that spans a visit. A visit-by-visit record answers what happened at the visit. A lifetime record answers questions nobody can ask of a visit: whether this patient's creatinine has been climbing for six years or stepped up last spring, whether the knee that hurts now is the knee imaged in 2019, whether a medication started for one condition is the reason a laboratory value drifted, whether the hypertension diagnosis was ever supported by readings taken outside a clinic room. Each of those is a trend question, and a trend cannot be reconstructed from documents written one at a time by institutions that never compared notes.
What arrives
| Kind of information | What it covers |
|---|---|
| Clinical records | Visit notes, diagnoses, procedures, hospital stays, prescriptions, signed forms |
| Testing | Blood work, pathology, imaging reports, every laboratory result over a lifetime |
| Imaging | Studies of any modality, stored once, with the signed report and measurements in the record |
| Genetics | Genome and exome data, epigenetic measures, telomere length, consumer DNA files |
| Family history | Conditions in parents, siblings and children, linked where people consent |
| Life history and exposures | Where a person lived and worked, and what was in the water and air there |
| Daily life | Diet, sleep, activity, supplements, wearable readings |
| Other practitioners | Reports from acupuncture, chiropractic, nutrition and other care outside conventional medicine |
| Device sessions | Treatment sessions, delivered settings and results from connected devices |
| What the person reports | Symptoms, how a treatment felt, check-ins three and fourteen days after anything new |
| Insurance claims | Claims history from Medicare Advantage, Medicaid and exchange plans |
| Outside records | Care summaries pulled from record networks, sent by secure message or faxed |
The original of every document is kept, write-once, and any value read from it links to the spot on the page it came from. Very large data with its own shape, such as years of per-second heart rate or a whole genome, is referenced from the event log and held in a store built for it.
The registry: a new kind of data without a database change
In a conventional system a new test, a new wearable metric or a genome means new tables or columns, a schema change with downtime, and edits to every program that reads them. The usual response is to push the new data into free-text notes or scanned attachments, where nobody can analyze it. A clinician who wants a number charted learns to type it into a note, and from that moment the number exists and is invisible.
Here a new kind of data enters through a registry entry and the database structure does not change. Every event has the same outer shape: who, what type, when, from where, how trusted, and a payload, and the registry defines what a valid payload for each type looks like. An entry states:
- What the type measures, in plain words.
- Its standard code where one exists: LOINC for laboratory tests, SNOMED CT for conditions, RxNorm for medications, DICOM for imaging.
- Its unit, and conversions from the other units it arrives in.
- The range of plausible values, so a typo or a misread photograph is held for review rather than landing.
- The method or device, when that changes what the number means.
- Its starting trust weight by source class.
- Its version, so a definition can be refined without changing what older events meant.
Adding a type takes a definition and a person's approval, and intake accepts the new data that day. A type with no standard code is registered locally and mapped when a standard one exists. Anything intake cannot map is kept and held for a reviewer, never dropped and never forced into the wrong field.
The types it ships with
Forty-six approved types cover what a deployment needs from its first day.
| Group | Types | Standard the registry maps to |
|---|---|---|
| Common laboratory tests (17) | Glucose, hemoglobin A1c, cholesterol, HDL, LDL, triglycerides, creatinine, sodium, potassium, hemoglobin, white cell count, platelets, TSH, ALT, AST, ferritin, TPO antibodies | LOINC |
| Vitals (8) | Systolic and diastolic blood pressure, heart rate, weight, height, body mass index, temperature, oxygen saturation | LOINC |
| Clinical facts (7) | Condition, medication, allergy, immunization, procedure, encounter, smoking status | SNOMED CT, RxNorm, CVX, LOINC |
| Imaging and documents (2) | Imaging study, document | DICOM; document types by LOINC |
| Devices and care (3) | Device session, care session, care task | Local codes |
| Patient-reported (6) | Patient report, patient question, meal log, sleep from a wearable, visit synopsis, photograph measurement | Local codes |
| Safety (1) | Triage flag, written when the emergency path raises a concern | Local code |
| Record keeping (2) | The chained record of a lawful erasure, and the type an erased row carries | System types |
A worked entry, for a hemoglobin A1c: the type maps to LOINC 4548-4, reports in percent, converts from millimoles per mole by a stated formula, holds anything outside 3.0 to 20.0 for review, and starts a certified laboratory at a weight of 0.95 and a value the patient typed at 0.6. Those two figures are the published illustration; a deployment's tuned weights, prior strengths, decay rate and declared source families are trade secrets and appear in no published file. Still ahead: full terminologies rather than mapped subsets, sensitivity labels, and specimen-aware mapping so a urine glucose can never map to a blood glucose.
Beside the record: graded knowledge, no patient data
Medical knowledge is kept in a separate store holding no patient data: papers, drug labels, guidelines, trial listings and literature on care outside conventional medicine, each with its source and retrieval date, and summaries tied sentence by sentence to that source. Every guide in the suite reads from it, and so does the evidence engine.
It grows from what patients raise, without carrying anyone's identity. A patient's question brings in a topic. If the library holds that topic and the entry is fresh, the guide answers at once. If the entry is missing or stale, a research job searches by topic words only, with identifiers stripped, and writes what it finds; the evidence engine then scores it. The next patient with the same concern gets the answer immediately, with a reference on every statement. Three stores stay separate: the patient record, under identity and consent controls; the knowledge library, read by every guide; and the scored claims formed from the library and from population data.
What a complete record makes possible
Four things a visit-by-visit record cannot support.
The trend, not the snapshot
A value read against six years of that person's own history is a different fact from the same value read alone. A result in range for the population and far from this person's own median is caught, and one consistent with the person's trajectory is not flagged for no reason. The record checks a new value against the person before the reference interval.
Outcomes the model can learn from
Learning which sources deserve trust requires knowing what happened afterward, and a record that ends at the visit never sees it. A lifetime record watches a drifting home cuff lose agreement with clinic readings week by week, and that cuff's weight for blood pressure falls below every well-behaved cuff on its own.
Questions answerable only over years
Whether a protocol worked, for whom, and for how long is a question about the years after treatment. Device outcomes, patient-reported response, laboratory trajectories and later events sit in one record, which is what lets a result be scored by indication and tissue rather than by whoever happened to follow up.
A history that survives changing doctors
Moving, changing insurers or changing physicians costs a patient their history in the conventional arrangement, because the history belonged to the institutions. Here it is the person's, and a new clinician inherits it rather than starting a chart.
A value with no history behind it is a number. A value read against six years of the same person is a fact.

HOLARC
Any format in
Any record a hospital system, laboratory, imaging center, device or person produces goes into Holarc as it is. Intake works out what it is, whose it is, and where each value belongs, and records where every value came from and how sure it is of what it read. If a person could read it, intake accepts it.
Why intake is the hardest part of a health record
Health records fail at the door. The parts that get built are the satisfying ones: a schema, a chart screen, an interface specification. The part that decides whether the system is worth anything is the part that takes in what the world actually sends, which is not what the specification describes. What the world sends is a faxed printout with a coffee ring on it, a phone photograph of a discharge summary taken at an angle, a spreadsheet export whose column headers changed last quarter, a laboratory message that differs from the standard in four undocumented places, and a care summary whose medication section is empty because the sender did not populate it. Every one of those is a normal Tuesday, and every one has to end up as a coded value with a unit, attached to the right person, with an honest statement of how confident the system is.
Most records never get the data in, because the alternative looks cheaper at every decision point. Requiring the sender to convert first is cheaper than reading what they sent. Dropping a file that will not parse is cheaper than keeping it. Guessing which patient a document belongs to is cheaper than a review queue. Together those choices produce the system everyone recognizes: a well-designed structure with almost nothing in the structured fields, and the real history in a folder of scans nobody can query. Holarc takes the opposite decision at every one of those points.
What it reads
| Source | Formats | State |
|---|---|---|
| Hospital and EMR systems | HL7 v2 messages and batch files; FHIR R4 in JSON and bulk form; care summary documents; vendor spreadsheet exports; chart printouts as PDF | Built, including the rule that a corrected, preliminary or withdrawn result updates one result rather than creating a second |
| Laboratories | HL7 v2 result messages, FHIR diagnostic reports, spreadsheet result files, PDF reports, faxes, photographs of printed results | Built |
| Paper and photographs | Scanned PDFs, faxes, phone photographs, word processor documents, plain text | Built, with orientation detection and text recognition |
| Imaging | Studies of any modality and size, zipped exports included, with their reports | Built: study details land in the record and the images go to the sealed imaging store |
| Devices and wearables | The native device feed for sessions and results, matched to the person at the point of care; standard FHIR resources; files from watches, rings, scales, cuffs and glucose monitors | Device and tabular readers built, with no device yet connected; phone platform exports recognized, archived and held |
| Genetics and claims | Sequence and variant formats, consumer DNA files, standard claim and remittance files | Recognized, archived and held; readers planned |
| Anything else | Spreadsheets, images, zipped folders of any of the above | Zipped folders unpacked within per-deployment limits. A zip that breaks a limit, names a path outside itself or holds a symbolic link is kept and held unread. Anything unrecognized is archived, fingerprinted and held as unreadable |
How a value becomes a fact
Every drop passes the same seven steps, and ends in a sealed report naming what landed, what waits and what failed.
- Receive. The original is archived first, sealed, and stored under the fingerprint of its bytes, with who sent it, when, and under what authority: a business associate agreement, a patient's authorization, the patient themselves, or a device during treatment.
- Recognize. A file is identified from its bytes rather than its name, and standard formats are parsed by fixed rules. Zipped folders are unpacked within their limits, each file handled on its own.
- Read. Structured files are parsed field by field. PDFs, faxes and photographs are read by text recognition, tables included, and every value is tied to the region of the page it came from.
- Map. Each value is matched to a registered type and a standard code. The first file from a new source gets a mapping a person approves once; that becomes the source's profile, and every later file from it maps on its own.
- Identify. Each record is matched to exactly one person, under rules that refuse to guess.
- Score. Each value gets two numbers: how sure intake is that it read the value correctly, and how far the record trusts whoever produced it for this kind of fact.
- Route. Values that pass enter the event log. Anything that cannot be matched to a person, fits no registered type, has no unit, falls outside its plausible range or below the confidence threshold goes to the holding bay. Each file runs in its own savepoint, so one corrupt file never rolls back the rest.
The checks catch the failures that matter rather than the ones that are easy to catch. A number must fall in its type's plausible range. A declared unit converts by the registry's formula, with the original kept beside the converted value, and a value below range that would make sense in another unit is held with that suggestion attached. Diastolic must be below systolic. A value from weak evidence sitting more than two and a half times above or below the person's own median is held, which catches a height of 67 entered in centimeters and a recognized creatinine of 10.9 that should read 1.09. A fasting glucose collected within eight hours after a logged meal is flagged and kept out of fasting trends, rather than quietly becoming a diabetic-range result.
A document that arrives as a photograph
A phone photograph of a laboratory report is a normal input. Intake detects its orientation, straightens and cleans the image, reads it and keeps it, and a value read from it carries the exact region of the image it came from, so the number and its evidence stay together.
When confidence is low, the value neither enters the record nor disappears. A reviewer sees it beside the crop it was read from and confirms or corrects it with one click, and every confirmation measures reading accuracy, so error rates become known by source and format rather than assumed. In the photograph test the recognizer read "Hemoglobin A1c" as "Hemoglobin Alc"; intake matched it to the registry with a confidence penalty, and a glucose line read at 0.81 confidence waited for a reviewer, as intended.
That path also produces the measurement a record needs about its own reading. When a laboratory feed reports a hemoglobin A1c of 5.8 and a photograph of the same report reads 5.3, both are kept, the laboratory outweighs the photograph, and the disagreement counts against the photograph reader's learned weight for laboratory results. The reader's accuracy is a running measurement the record takes of itself rather than a vendor claim.
An AI document reader for handwriting and unusual layouts is built and ships switched off. It turns on after its error rate is measured against reviewer-confirmed values, and it reaches a model only through the suite's gateway, which strips identifiers, checks consent and logs the call.
The holding bay, and the human review path
Nothing that arrives is lost. Records intake cannot place wait in the holding bay with their original file, the reason they are waiting and a suggested answer, all sealed at rest. A steward clears the bay, each item is decided once under a lock, and every decision is a row in an append-only table under that steward's own token. That approval path is the difference between a record that accumulates data and one that accumulates data it can use: the first file from a new laboratory takes a person ten minutes, the ten thousandth takes nobody any time, and every value in it carries the provenance the first established.
What intake never does
- It never guesses a person's identity when the match is uncertain.
- It never changes an original file or creates a value the source did not contain.
- It never discards a file because it could not read it.
- It never accepts data without recording who sent it and under what authority.
Every original is kept unchanged, stored by its fingerprint and sealed, so any number can be checked against its source years later. An original is deleted only by lawful erasure, and only when nothing else refers to the same bytes.
If a person could read it, intake accepts it. No sender converts anything first.

HOLARC
One record per person, and who may see it
Each person gets exactly one record. Two failures threaten that: a duplicate that splits one person across two records, and a false merge that puts two people into one. The design guards against the false merge first, because it puts one person's allergies and medications into another person's chart.
The identity vault
Every person receives an identifier that is a random number and means nothing outside the record. It is never a Social Security number and never built from one. Names, birth dates, addresses, phone numbers and every outside identifier live only in the identity vault, and the clinical event store sees a random token. A breach of the records exposes no names, and a breach of the vault exposes no records.
Each field in the vault is sealed on its own, bound to its row. Exact-value searches go through blind indexes, meaning keyed hashes under a separate key purpose, so the ciphertext is never searched and a copy of the database taken without the keys opens nothing.
How an incoming record finds its person
Most identity confusion comes from outside identifiers rather than names. A clinic knows a person by its medical record number, an insurer by a member number, a device by a serial number. The vault keeps a crosswalk from each of those to the one record, so once a clinic's number is linked, every later record from it matches exactly.
- Exact match first. A record carrying an identifier already in the crosswalk goes to that person.
- Identified at the point of care. At a device or a clinic the patient presents a code from their own application, or staff look them up and confirm, so device results never need a guess.
- Scored match for everything else. Records with no known identifier are scored on name similarity, birth date, sex, address, postal code, phone and email, allowing for spelling variants, nicknames and moves. This is standard probabilistic matching, used in hospital patient indexes for decades.
- Three outcomes. A high score links automatically. A low score starts a new record. Anything in between goes to a reviewer and, where possible, to the patient, whose application asks whether the visit was theirs.
Some rules no score can override, because the cost of the two errors is not symmetric. A match on name, birth date and sex alone never links automatically; a phone, email, address or postal code must also agree. A different given name, a different sex, a suffix conflict, a near birth date, or a close second candidate all send the record to review. Identifiers from sources that cannot be named are never used for matching.
Details typed into an application never link automatically however strong the match, and a steward's decision that two of them are the same person records the link and nothing else. Access to the record on file is a separate step requiring proof of control.
The cases that break matching
| Case | How it is handled |
|---|---|
| Newborns under a temporary name | The child's record is created linked to the mother's at birth, so the link never depends on the name |
| Twins, juniors and seniors | More agreement is required before linking, and these go to review. Twin pairs and two men of the same name born the same day are in the tests; neither links |
| Name changes and moves | The vault keeps every past name and address, so an old record still matches |
| A person who is erased | Erasure covers the canonical person and every record merged into them |
Because the record only adds, a merge is a recorded entry and so is an unmerge, which separates two records joined in error exactly and lets every view rebuild correctly. Only a steward may do either, and the automatic-link threshold is set so doubtful cases go to review, trading a few more duplicates for fewer false merges.
On the invented 519-person population the matcher scored 62,296 candidate pairs, linked 2,229 of the 2,257 true same-person pairs automatically, sent 14 pairs to review, and made no false merge. Matching accuracy on real records is a publication owed before sale, because an invented population is not evidence about real names.
Verified identity, and who may act for whom
The design anchors every record to an identity verified to the federal standard for assurance level 2, which the federal record-network program requires before a patient application may pull records. Today patients sign in with a one-time code sent to the contact on file, staff with a hardware security key, and proxy access runs through a proofing adapter on a simulator until a service is contracted. Proofing for every patient at sign-up is planned before real patients.
A proxy, meaning a parent for a minor child, an adult granting access, or a legal representative, is a recorded link between two records, each with its own identifier, and one person's data never moves into another's. Proxy access for a minor ends at eighteen, and records a minor may lawfully keep from a parent are withheld. Those limits differ by state and by the kind of care, so they sit in the consent engine rather than in an application.
Two custodians over one set of events
A clinic is the legal custodian of its chart. It must keep it, seven years for adults in California and at least a year past a minor's eighteenth birthday, and take all of it when it leaves. The patient owns the personal record, and a copy of each finished clinic record flows into it under the patient's federal right to direct copies to an application of their choice.
| Store | Custodian | Patient may erase |
|---|---|---|
| Clinic record domain: notes, orders, results, medications, allergies, problems, image links, signed forms | The clinic | No. State law makes the clinic keep it |
| Personal record domain: the patient's copy, plus everything they and their devices add | The patient | Yes |
| Clinic operational store and patient accounts: schedules, drafts, charges, claims, denials, balances | The clinic, under its own keys | No |
| Practice ledger: deposits, expenses, payroll, taxes, no patient identity | The clinic, under its own keys | Not applicable |
The personal domain can be erased without touching the chart. Clinical codes flow into a claim one way when it is built and are never copied back, and nothing in a clinic's accounts or books has a query path to the record, to research or to another clinic. A clinic may hold its own master key, and operator access is logged for its administrator.
What a clinic sees against what a patient sees
A clinic sees its own chart and the outside records the patient approved that clinic to see. Its staff token reaches only the people on that clinic's rosters, on every route alike, so a clinic cannot browse the deployment. A patient sees their whole record, every consent they signed, their dependents under a recorded legal basis, and a log of every read, allowed or refused, naming who looked, why, and which application they used. Refusals are logged as well as permissions, which makes that log worth showing: a patient who can see who read their record can tell whether their consent is being honored.
How a licensed partition is separated
Another company can operate a partition inside the same record: its own data, its own keys, its own trust settings and its own views. The walls are the ones already enforced between custody domains rather than new code written for the licensee, which is the only reason they can be trusted. A wall that exists as an application setting is a wall somebody can change. Data never crosses one on its own: a person in a licensee's partition who also holds a personal record has the two linked only through the identity vault, only when they authorize it, and the link is a logged event.
A duplicate is an inconvenience. A false merge is a clinical event.

HOLARC
Connecting to the systems that already hold the data
A lifetime record is only as complete as what reaches it, and almost everything that should reach it is already sitting in somebody else's system. Holarc owns the exchange interfaces for the whole suite: one implementation of each standard, used by every product, rather than each product writing its own.
A record that cannot connect is worth nothing
The data a person needs in one place is produced in a hundred places, and none of them will send it anywhere new. A hospital will not build a custom feed. A reference laboratory will not change its message format. A regional exchange will not add an interface for a record system with no customers. What they will all do is answer the standard interfaces they are already obliged to answer, because federal rules require certified systems to offer them.
That obligation is the whole opening. A patient can direct any certified system to release records to an application of their choice, and the information-blocking rules make that a request the system must honor. A health plan must offer a patient access interface. A record network answers a query for a person's documents. A laboratory sends results in a format defined thirty years ago that every laboratory still speaks. A record implementing all of that arrives at a new clinic already able to pull the patient's history; one that does not arrives asking the clinic to type it in, and loses.
The cost is why most records skip it. Each interface is a specification of hundreds of pages, a conformance test suite, a partner contract, and a long tail of other people's implementation quirks. It is also the part a competitor cannot route around.
What each interface is, in one sentence
| Interface | What it is | State |
|---|---|---|
| FHIR R4 with US Core | The modern web interface for health data: a request for a patient's conditions, medications, laboratory results, immunizations, procedures, encounters or documents returns them in a shape every certified system agrees on. | Built and served by Holarc |
| SMART on FHIR | The sign-in standard in front of it, so a patient or clinician authorizes an application to read a named scope of their record and nothing else, and the application never holds a password. | Built, with the suite's own authorization server |
| Bulk data export | The same interface asked for a population rather than a person: a clinic's whole roster exported for analysis, quality reporting or migration. | Built |
| C-CDA | The document standard that carries a patient summary between institutions at a transition of care: allergies, medications, problems, results, vitals, immunizations, procedures, encounters and social history in one signed document. | Built, generated and imported |
| Single-patient export | Everything held about one person, packaged with a manifest, which federal rules require a certified system to be able to hand over. | Built |
| Record networks | The national and regional networks that answer who holds records for this person and then send those documents, across participating institutions under a common agreement. | Built against a simulator; a partner is contracted at deployment |
| Direct secure messaging | Encrypted clinician-to-clinician email on a trusted certificate network, which is how a referral or a discharge summary travels when there is no shared system. | Built against a simulator; a carrier is contracted at deployment |
| Fax | Still how a large share of American medical records actually move, so inbound faxes arrive as documents and go through intake like any other file. | Built against a simulator; a carrier is contracted at deployment |
| Immunization registries | The state registries that hold immunization histories and expect a specific message format with its own acknowledgment codes. | Built against a simulator; each state requires its own onboarding |
| Public health case reporting | The report a clinician owes a public health authority when a reportable condition appears on a problem list. | Prototype; messages are produced and kept, not sent |
| Quality measures | The federal clinical quality measures a practice reports on, computed from the record rather than from a separate registry a practice feeds by hand. | Prototype; certification needs a certified measure engine |
The standards interface in detail
Read and search are served for patients, allergies, conditions, medication requests, observations covering laboratory results and vital signs, immunizations, procedures, encounters, document references with the original behind them, and diagnostic reports. Search parameters follow the required server behavior for each resource type. A patient resource carries the canonical record identifier even after merges, every other resource carries the identifier of the event it was made from, and corrected values are served as the correction. The interface is read-only, which is deliberate: data enters through intake, where it is identified, mapped, scored and recorded with its provenance, and a write interface would be a way around all of that.
Every resource Holarc serves carries an evidence extension with five parts: the source, that source's learned weight now, the weight it carried when the value was written, the extraction confidence, and the method. A blood pressure panel carries it on each component. That extension makes the evidence-weighted record visible to any standards-compliant client without that client knowing anything about Holarc, and it is how an outside application gets the one thing no other record can give it.
Authorization is enforced three ways at once. A scope grants read and search on a named resource type for a named subject class. A patient's token is always held to that patient's own record, so a token naming someone else grants no context rather than the wrong context. A clinic's token reaches only the people on its rosters. Every read and refusal goes to the person's access log, naming the user and the application. Bulk exports fix their population at kick-off, check consent for each person as the job runs, name anyone whose consent no longer covers the read, and seal and expire the output files.
Inbound: everything becomes a document, and a document becomes evidence
Records arriving from outside never enter directly, whatever route they came by. A summary pulled from a record network, a document sent by secure message, a faxed printout and a file a patient uploads all take the same intake path: read, mapped to registered types, matched to one person, scored, and either entered with its provenance or held for a reviewer. The sending institution becomes a source like any other, with a weight learned from outcomes rather than assumed from its size.
Nothing from secure messaging or fax enters the record until a steward files it, and each inbound document is decided once. Senders, fax numbers, subjects and attachment names are sealed, and the attachments go to the sealed archive. A network query records its result as complete, partial, no match or error, with identifiers and counts only, and one community failing never stops the others. Where a network returns more than one candidate person, nothing is pulled from that community at all.
What runs against a simulator today, stated plainly
Every outside network sits behind one adapter. The simulator is the default, and it behaves like the real network, rejections and error codes included: synthetic communities with a document to find, a registry that fails a query, and a network whose answer arrives in a different format. The real adapter for each network exists and refuses to run until the thing it waits on is in place, naming what that is. A deployment switches one network at a time, after that partner's contract and onboarding.
What that means for a reader evaluating this: the interfaces Holarc serves, meaning the standards interface, the document standard and the two exports, are built and need no partner. The interfaces Holarc consumes, meaning the record networks, secure messaging, fax and the immunization registries, are built and have never spoken to a real counterparty. The messages they produce are the real shapes, down to the envelopes and acknowledgment segments, and the simulator rejects malformed ones. That is not the same as having exchanged a record with a real institution, and the difference is a contract and a conformance test rather than more code.
Four things wait on something other than code. A consent screen for patients authorizing an outside application. A network partner, a secure messaging carrier, a fax carrier and state registry onboarding. Federal certification testing against a deployed instance, including the document validator and the conformance kit. And a certified measure engine, which cannot be substituted.
The standards are the standards. A record that implements them arrives able to pull a patient's history; a record that does not arrives asking a clinic to type it in.

HOLARC
Tamper evidence and lawful erasure
A lifetime record is relied on by people who were not there when it was written: a physician reading a laboratory history, an auditor, a court years later. It has to prove nothing in it changed after it was written, without exposing anyone else's data, and keep proving that after the law requires part of it to be deleted.
Two requirements that usually cancel each other
The ordinary way to prove a record was not altered is to make it immutable. The ordinary way to satisfy a privacy law is to delete what a person asks to have deleted. Those instructions point in opposite directions, and most systems abandon one: a record on immutable distributed storage cannot honor an erasure request, and a record built to delete freely cannot prove anything about its own history. Holarc holds both, and the decision that makes it possible is to commit to each value through a salted digest rather than the value itself. The chain links survive the removal of the content they were computed over, so a person's data can be erased while every proof issued before the erasure stays valid.
That is also the reason a blockchain is not used. A blockchain guarantees an entry was not altered after it was committed, which Holarc needs, and says nothing about whether the entry was true when it was committed, which is the question an evidence-weighted record exists to answer. Immutable distributed storage also works against the erasure state privacy laws and European law require.
How the chain works
Each event carries a cryptographic fingerprint of itself and of the event before it. The content fingerprint covers a canonical encoding of the body together with thirty-two random bytes generated fresh for that event, and the link covers the previous event's link and that fingerprint. A link can therefore be checked without the body, which lets the chain survive erasure, and any change to a past event breaks every link after it. The fingerprint covers the plaintext before the value is sealed, so rotating keys changes no fingerprint, while a ciphertext swapped between rows fails its authenticated decryption and shows as a broken link.
The random bytes are not decoration. A birth date, a postal code or a diagnosis code has few enough possible values that anyone holding a plain hash of one could test them all; with random bytes mixed in, and those bytes deleted at erasure, that test is impossible. The hash function is the one national cryptographic policy requires for long-lived data.
The head is anchored outside the database
A chain living entirely inside a database defends against a careless administrator and not a determined one. Someone with full control can rewrite an event and recompute every fingerprint after it, and the result is internally consistent. What catches that is a copy of the chain's position held where the database owner cannot reach.
Every hour the record verifies its chain locally and posts four numbers to the anchoring service: the height, the head fingerprint at that height, a Merkle root over every event fingerprint, and the tree size. The service refuses another product's credentials for this namespace, a tree smaller than the last it holds, a second root for a size it holds, and a different head at a height it holds. Each refusal is a rewrite caught at the door.
It signs each anchor twice, with a classical signature and with the standardized quantum-resistant signature, and both must hold. Anchors are chained to each other and written to an append-only ledger file on a separate volume, which the nightly encrypted backup carries off the host, and each can be timestamped by independent authorities so its date rests on no single party. Anyone holding that file and the two published keys can check every anchor offline, with no database, network or secret. Moving the signing keys inside hardware waits on hardware support for the quantum-resistant algorithm.
A consistent rewrite of the most recent event leaves every link and the stored head in agreement, so the database alone cannot catch it: the local verification passes and the anchor check fails. Both outcomes are asserted in the tests, so nobody reads a green local verification as evidence of a sound record.
Proving one record without exposing any other
An auditor, a court or an insurer asks about one record, and showing them the chain would show them everyone. A single-record proof answers without that exposure. It holds the event's fingerprint, its content fingerprint, the previous link, the Merkle audit path up to the newest anchor covering it, roughly twenty fingerprints for a record of a million events, and the signed anchor, and it can carry that record's body and random bytes decrypted so the reader checks the content as well as its position. No other record's content, identifier or fingerprint appears in it. One offline script checks a proof from either record in the suite, and both implementations produce identical bytes for the same input, which makes a proof portable rather than a thing only the issuer can read. A record since erased can still be proved to have been present and never disclosed, which answers what an auditor asks about a deleted record: whether the deletion was a deletion or a cover-up.
Lawful erasure
An erasure runs under a stated legal basis, with a reference and an actor recorded, and covers every record resolving to the same canonical person. In one transaction the system marks the person erased, blanks each of their events including values, text, units, dates of service, source and authority, deletes the random bytes, writes one chained erasure event naming the rows and the basis, clears the vault, and removes values of theirs still in the holding bay. Each blanked row keeps its sequence, identifiers, times, content fingerprint and link, so the chain still runs through it. Original documents are deleted after the transaction commits, and only those nothing else refers to. Every earlier anchor and proof stays valid, and a proof made after an erasure can show the row was there and never what it said, because the random bytes are gone and no guess about an erased value can be tested against its fingerprint.
Verification is specific about the ways an insider could try to erase quietly. A run fails on a changed value, a recomputed fingerprint, a broken link, or a head or Merkle root differing from an anchor. It fails on an erased row that is not fully blank, or that no chained erasure record written after it names, or on an erasure record listing rows written after itself or leaving a live row of its own token unmarked, so no single row can be erased on its own. The database permits the erasure update only inside a transaction that declares itself, and only one that blanks content and keeps every chain column.
What a chain proves, and what it does not
A hash chain proves that what was written has not been altered, and says nothing whatever about whether it was true. A chain will attest with perfect fidelity to a laboratory result that was wrong, a diagnosis that was mistaken and a reading from a device drifting for eight months. Integrity and accuracy are different properties, and conflating them is the standard error in how this kind of system is sold. Accuracy is what the trust model answers, which is why the two mechanisms sit side by side: the chain fixes what was written, and the weights say how far each writer has earned belief.
Four further limits, stated rather than left to be discovered:
- It proves a change happened. It does not stop someone with full control of the database from making one; it guarantees they cannot hide it once the record was anchored.
- Events written since the last anchor are protected only by the chain inside the database, a window of at most an hour.
- Erasure removes content from the live record at once. Copies already in encrypted backups stay readable to whoever holds the keys until those backups age out. Per-person keys, which would make destroying one key render every copy unreadable at the moment of erasure, are designed and not yet built, and are what closes this gap.
- Images in the imaging store are counted in an erasure result and not yet removed by it. Public health submissions and the accounting of disclosures survive an erasure, because the law requires them kept.
A chain proves that what was written was not altered. It says nothing about whether what was written was true.

HOLARC
The Holarc market
Holarc sells access to a platform and never sells records. Its customers are the device companies, clinics, application companies and research sponsors that pay to work on it. The figures below are published estimates and institution counts rather than a commissioned study.
The market Holarc sells into
Electronic health records, worldwide.
Where the value moved
Retrieving a person's records used to be the hard part, and it is becoming cheap. The national exchange framework passed one billion records exchanged in June 2026 across eleven networks. The federal aligned-network program, which requires standard interfaces and verified digital identity, began going live between April and September 2026, and a patient's right to direct any certified system to release records to an application of their choice is enforced by the information-blocking rules. When retrieval is commoditized, the value moves to whoever holds original data from the source, keeps it continuous for the person, and can say how far each piece can be trusted.
Clinics are the bottleneck rather than patients: by April 2026 more than 700 companies had pledged to the federal program and fewer than 50 provider organizations had. Holarc reaches patients through clinic software and devices, the side of that gap with less competition in it.
Demand for what a record enables is not in question. ChatGPT receives about 300 million health questions a week, and Microsoft more than 50 million a day, all from people asking medical questions of a system that knows nothing about them.
Sizing, and what the published numbers are worth
Published estimates for the personal health record software market alone run from under $50 million to over $6 billion depending on where the definition is drawn, so the plan rests on none of it. The segments Holarc sells into directly are larger and better measured.
| Segment | Global, per year | United States |
|---|---|---|
| Electronic health records | $30.3B to $37.5B | $9.4B to $15.0B |
| Patient engagement and personal health records | $33.8B to $41.2B | $14.6B |
| Real-world evidence for research | $2.6B to $6.2B | About $1.0B |
| Healthcare analytics, adjacent | $78.3B to $81.9B | $28.2B |
Research firms define these markets differently, so each line is a range between the low and high published estimate, several overlap, and adding them together would count some spending twice. Counts of institutions are harder to argue with than market sizes built on other market sizes.
Those counts come from Definitive Healthcare, the federal payment advisory commission's March 2026 report, DPC Frontier, and the GCC Statistical Centre.
What investors pay for one layer of this
A better guide is what buyers have paid for companies owning one piece of what Holarc combines.
| Company | Layer it owns | Value | Date |
|---|---|---|---|
| OpenEvidence | Evidence answers for clinicians | $12B valuation | January 2026 |
| Commure | Clinic AI, scribe, revenue cycle | $7B valuation | May 2026 |
| Hippocratic AI | Patient-facing voice agents | $3.5B valuation | November 2025 |
| Innovaccer | Health data platform | About $3.45B valuation | January 2025 |
| Function Health | Consumer laboratories, longitudinal laboratory data | $2.5B valuation, then $450M more | November 2025, July 2026 |
| Truveta | Research data | Above $1B valuation | January 2025 |
None of those six owns more than one layer. The record under them is what each has to buy, build or do without.
Who pays, and for what
Holarc prices the platform and never the records. Federal privacy law treats payment in exchange for health information as a sale requiring each patient's written authorization, so there is never a per-record price. That constraint is also the strongest commercial position available: a counterparty worried about where its patients' data will end up is relying on a rule the vendor cannot profitably break.
| License | Who pays | What they receive | Structure |
|---|---|---|---|
| Device access | Device and diagnostic companies | Devices connected to the record; settings, completions and outcomes per patient; outcome scoring by protocol, indication and tissue; follow-up through the patient application; compliant storage, identity, consent and audit logs they do not build or certify | Per connected device per month, plus a fee per recorded treatment |
| Application access | Clinic groups, employers, health systems, wellness companies, developers | A branded edition of the patient application, or consented access through the standard interfaces, reading only the data their own application created | Per active patient per month; per connection for outside applications |
| Research project | Device companies, drug developers, universities, health systems | A defined question answered inside the walls on consented, de-identified data, negative results scored the same as positive ones | Per project, by cohort size and scope |
The device license is the one with no competitor. An aggregator receives what another system chose to send, usually a summary written after the visit. A connected device writes its own settings, its completion report and each patient response where they are created, into a record already holding that patient's laboratory results, imaging, medications and history. The company ends with scored evidence across every patient it treated, produced under rules it does not control, which is worth more to a regulator than evidence it graded itself. What makes that acceptable is the rule on methods: algorithms, settings logic, treatment curves and reconstruction methods stay on the company's own servers, because the record stores the result of using a method and never the method. The company keeps its trade secret and gains a long-term outcome record without building a record system or a compliance program.
The licensed-partition business
Another company can license a partition of the same record: its own data, keys, trust settings, registry additions and views, inside the walls the record already enforces. A company producing health data about people has a problem unrelated to its product. It needs somewhere for that data to live that will satisfy a hospital's security review, a regulator's audit, a buyer's diligence and a plaintiff's lawyer, for as long as the data must be kept, which in health is decades. Building that is a program rather than a feature, none of it differentiates the company, and all of it has to be right. The model assumes $250,000 a year for an average licensee, an assumption rather than a quoted price and deliberately the least load-bearing line in the business at 1.4 percent of modeled 2031 revenue. The strategic value is larger, because every licensee is a source of the independent outcomes the trust model cannot manufacture for itself.
The research answers business
Researchers bring questions. The analysis runs inside the walls on consented, de-identified data, and the answer comes out with every result checked before release so no small group can be singled out. Records stay where they are, which keeps patients, clinics and regulators on the same side of the question and means the research business is not traded off against the privacy position. The model assumes $150,000 for an average engagement. Null results are flagged, kept and earn the same credit as positive ones, and conventional care, alternative care and the suite's own affiliated products are scored by the same rules, which protects the record's standing as an honest broker and the affiliated products from the charge that their evidence was graded on a curve.
Who else is trying, and what they do not have
The aggregators are competitors and likely suppliers at the same time: the plan buys retrieval from them and competes above it.
| Competitor | Same as Holarc | Different |
|---|---|---|
| b.well | A longitudinal standards-based record, white-labeled for payers and employers | Built from pulled documents; no writing at the point of care, no trust weights, no imaging |
| Particle Health, HealthEx | Deduplicated or patient-authorized access across national networks | Retrieval at query time, no lifetime record. One lost its access to the largest incumbent in 2024 |
| Zus Health, Metriport | An aggregated, normalized patient profile | Embedded in other records; no published matching accuracy |
| openEHR and its implementations | New data types with no schema change, every change versioned | No learned trust weight, patient-owned custody or evidence engine. The closest technical relative |
| Verily Me | A consumer record with an assistant and a research registry | No clinic write-back or trust weights found. The closest architectural match |
| Epic | A continuous record inside its own customers, plus a research warehouse | Stops at its customers, leaving out independent practices, care outside conventional medicine, device data and the patient's own entries |
| Truveta | Longitudinal de-identified research data | Research licensing only; no benefit returned to the patient |
Among the competitors reviewed, none weighs each value by a learned reliability for its source, and none writes into the patient's record at the point of care and measures device outcomes against it. Candidate patent families cover the learned trust weight per source and kind of fact, the fading memory, declared source families, the device outcome loop and split custody in one record.
Retrieval produces documents. A record produces a history.
MY TALISMAN
My Talisman
My Talisman is the patient's own app and voice guide, and the front door to their lifetime health record. It gathers records from anywhere, tracks medications and supplements, prepares patients for visits, answers medical questions with graded evidence, and carries the whole family. It informs. It never diagnoses and never prescribes.
The patient is the most complete source in health care, and the scarcest one
A person's health history is produced by institutions that never meet. A primary care office, three specialists, two hospitals, a laboratory, an imaging center and a pharmacy chain each hold a piece of it. Only one party was present at every encounter, knows which supplements are in the cabinet, and can say whether the treatment worked. That party is the last one anybody asks.
Giving the patient a guide over a lifetime record turns that around. A record that reads any format, carries the whole family, pulls from the national networks and connects to any clinic makes the patient the most complete source in the system rather than the least reliable one. Claims data arrives late and describes billing. A clinic note describes what the clinician concluded, which is the system's own answer coming back to it. A patient check-in on day 3 and day 14 after a new medication describes what happened to the person, an independent judgment the trust model under the record cannot manufacture for itself.
What the app does
The patient talks or types to a guide that already knows their whole history and does the work of staying healthy with them.
Every record, from anywhere
From the national networks, the patient-access interfaces certified systems must offer, insurer claims, uploads, photographs of paper reports and connected devices. Imaging arrives as full studies rather than a report about one.
What nobody has connected yet
A morning brief reports what the guide noticed: a ferritin falling across three draws, a supplement added Monday that affects a thyroid prescription, a screening due, tomorrow's visit with the questions drafted.
The visit, and what came of it
A prep brief the night before, notes taken with no recording kept and shown to the clinician for approval, then a plain-language summary with tasks, reminders and timed check-ins.
Any medical question
Conditions, tests, drugs, supplements, herbs, diets, acupuncture and the rest, each answer carrying a reference on every statement, a grade on every source and a score capped when the research is thin.
The whole family
Children, a spouse who granted access, a parent under a power of attorney. Each keeps a separate record under a link storing its legal basis, its signature and its end date.
A clinician in one tap
Telemedicine with a visit packet built from the record and approved by the patient first, their own practice offered ahead of anyone else. Sharing with any outside doctor through a revocable link.
Medication tracking covers prescriptions and supplements together, and a new diagnosis triggers research before the patient asks: a plain overview, treatments with their evidence strength, and recruiting trials filtered for that patient.
Sign-up needs no app
A clinic sends an invitation by text, email or QR code, and it opens in any browser. The patient signs in with a one-time code sent to the phone or email the clinic already has on file, and the guide walks them through demographics, insurance card photographs, intake forms, consents and signatures as they go. The account then exists against one record identity, and installing the phone app later signs into the same account with everything already in it.
That removes the step where patient engagement products lose most of their users. A front desk that can send a text message can enroll a patient, and the guide finishes it rather than the staff.
It never depends on the clinic
My Talisman requires no portal connection and no cooperation from the patient's doctors. The record is assembled from the networks, the certified patient interfaces, claims, uploads and devices whether or not a provider in it has heard of the company. Patients of practices running competing systems are served identically.
When a patient's clinic does run Arca, the clinic functions switch on inside the same app under the clinic's own logo: booking, intake forms and consents, released results, secure messages, bills, refills and video visits, plus access for a parent or caregiver after identity proofing. The clinic gets a modern portal it did not build and the patient gets one app instead of five. Arca never requires a patient to install anything, and My Talisman never requires a clinic to buy anything.
Wellness plus, and the line it does not cross
My Talisman is a wellness product with more reach than wellness products have had. It explains conditions, research and treatment options including experimental ones, tells the patient what to ask their doctor, flags interactions, and says when a symptom means going to an emergency room now. It never diagnoses, never rules a condition out, never prescribes or recommends anything beyond over-the-counter products, and never tells a patient to start, stop, skip or re-time a prescription.
| The guide does | The guide never does |
|---|---|
| Explains conditions, research and treatment options | Diagnoses a condition, or rules one out |
| Tells the patient what to ask their doctor | Prescribes, or recommends beyond over-the-counter products |
| Flags an interaction and sends the patient to the doctor or pharmacist | Tells a patient to stop, start or change a prescribed dose |
| Suggests over-the-counter products with label directions | Says a wearable reading means a specific disease |
| Says when a symptom means calling 911, and routes crisis wording to 988 | Presents itself as a person, or judges a photograph of a skin spot |
The line is held by two layers rather than by a prompt. Emergency rules read what the patient said before anything else happens, and a model layer may raise the urgency and can never lower it. An answer filter then reviews every answer, spoken or typed, before the patient sees or hears it, and replaces any sentence that crosses the line. The filter reads each sentence for who it is about, whether it hands the question back to the doctor, and what a negation covers, so "don't stop taking it without asking your doctor" passes and "don't take your metformin tonight" does not.
Every feature is classified against the federal general wellness and clinical decision support guidance before it is built, and is designed so it can move into a cleared path later rather than be rebuilt for one. The regulated claims in this business belong to licensed people and cleared devices. The guide earns its place by being the best-informed non-clinician in the room.
Who pays
Patients never pay to own their data, and records are never sold, rented or transferred. Organizations pay for the app, the guide and the record behind them.
| Buyer | What they get | Structure |
|---|---|---|
| Clinic groups, employers, health systems, wellness companies | A branded edition from one code base, set by a brand pack: logo, colors, welcome message, care team, feature switches, and usage reporting with nobody identified | Per active patient per month |
| Arca clinics | My Talisman as the clinic's patient portal, at no cost to its patients | Included in the Arca subscription, priced per clinician per month |
| Outside application developers | Consented access through the standard interfaces, reading only what their own application created | Per connection |
| Individuals | The direct edition under the My Talisman name and colors | Price undecided; the terms require a shown price and agreement before any charge |
Prices are set with the first partners during the pilot. My Talisman stays one product on one code base, so every branded edition carries the next improvement to all of them.
Consumer health apps win patients one at a time. A clinic brings its whole panel on the day it signs.
MY TALISMAN
The guide, by voice or text
Every patient meets a guide with its own name and voice, one that knows their whole record and never asks them to repeat their history. Anything the patient can say they can type, and the typing window never covers the screen.
A guide the patient names
During intake the guide suggests a name suited to that patient and offers alternatives, so two patients rarely meet the same guide. The sample patient's guide suggests Wren. The patient can take it, pick an alternative, type a name or say one out loud, and can rename it at any time. They choose a male or female voice at setup and can switch whenever they like.
The naming is not decoration. A health product's hardest problem is that nobody opens it when they are well, and the thing people return to is a relationship rather than a dashboard. A guide with a name the patient chose, a voice they picked and a memory of every conversation is something a person talks to. The brand pack carries the organization's logo and colors; the guide's name and personality stay the patient's, so a clinic that changes software does not take their guide away from them.
When asked what it is, the guide says it is an AI. It never presents itself as a person or a doctor, in any edition, under any brand.
Three habits
- It knows without being told. The patient never repeats their history, never re-enters a medication list, and never explains which specialist ordered which test.
- It has done the work before it is asked. The morning brief connects things nobody has connected yet: a ferritin falling across three draws, a supplement added Monday that interacts with a thyroid prescription, a child's immunization due.
- It speaks up at the right moment. A check-in on day 3 and day 14 after a new medication, a screening that is due, a visit tomorrow with the questions ready. Those check-in answers are recorded as outcome data, the independent judgment the record learns from.
Talk or type, with nothing lost either way
Nobody should have to talk about their health out loud on a bus, or have a phone say it back to them in a waiting room. Every patient chooses at setup how to use the guide and can change it at any moment. The choice follows them to every device.
| Mode | How it behaves |
|---|---|
| Out loud | The patient talks and the guide answers out loud. While the patient is talking the typing panel closes so the whole screen stays in view. The guide opens the screens it is talking about, and a slim caption at the top carries its latest words. Spoken answers stay short because the full answer is on the screen. |
| Typing only | The guide never speaks and never asks for the microphone. The guide button opens the typed conversation and the dictation controls are hidden. At a visit, the notetaker shows its permission request on screen for the clinician to read. |
| Desktop | The conversation sits in its own column beside the screen and never covers it. |
Every voice function has a text equivalent. No function in My Talisman requires speech or hearing.
Anything the patient can say, they can type. Anything they can enter, they can say. Every screen can be driven by voice, and the guide shows what it heard before it saves anything. Switching is one sentence: saying "type instead" ends the speech on the spot. That rule is a design constraint rather than an accessibility afterthought, and it decides product questions on its own, since a feature that would only work spoken does not ship. It also settles the two objections that follow voice products everywhere, privacy in public and users who will not talk to software.
Teaching by doing
The guide does not hand the patient a help article. For any task, uploading a laboratory report, connecting a watch, adding a medication, sharing a synopsis, it explains the step in one sentence, highlights the next control, and waits while the patient does it. It will do the step for them when asked.
- Stuck detection. Repeated taps in one place, or a long pause on a screen, prompt the guide to offer help before the patient gives up.
- Pace. Practiced patients get shorter walkthroughs, and a simple mode slows down, uses larger type and shows one step per screen.
- Caregivers. A family member managing someone else's record gets the same step-by-step teaching as the patient.
- Onboarding is itself a walkthrough. What the guide should be like, what to call it, talk or type, the voice, then each agreement in plain language. The patient signs each one; the guide never accepts for them.
The commercial value is the training line. Patient engagement products are sold to organizations that will not run a training program for their patients and cannot staff a help desk. A product the guide teaches, function by function, has no training cost to pass along.
One guide service, two sets of tools
The patient's guide and the clinic's guide inside Arca are one guide service with different tools, permissions and voices. A clinician's guide can order, chart, bill and schedule under that clinician's authority. A patient's guide can read that patient's own record, answer questions with graded evidence, open screens, run a walkthrough, take visit notes, request a video visit, share the record, rename itself and switch to a family member. Neither can reach the other's tools.
Building it once has three consequences. Every improvement to speech handling, the guardrail layer, walkthrough teaching or the conversation model arrives on both sides at once. The safety work that matters most, emergency triage and the answer filter, is written once and tested once rather than twice with two sets of gaps. And a patient and their clinician talk to systems that understand the same record the same way, which makes a clinic portal inside a patient's own app coherent rather than bolted on.
What stays hidden
Patients never see the names Holarc, NERD or Keelson, or any other system name. They see their guide, their record and their clinic's brand. A test checks that no system name reaches a patient screen; the names appear only in the license terms, where ownership of the technology is stated.
The architecture underneath, a lifetime event record, a learned trust weight per source, an evidence engine, is what makes the guide better than a chat window over a document pile. None of it is the patient's problem. A patient asked to understand a record engine has been handed the vendor's organizational chart instead of an answer.
How a spoken turn is handled
Every voice conversation starts on the server rather than in the browser. The service asks the suite AI gateway for a signed session; the gateway holds the vendor credential, checks the person's voice consent and sets recording off with minimum retention. The browser receives a signed address, a brief of the record and a one-time bridge token, and never holds a vendor key.
- Triage first. The voice agent's language model is the service itself. Each turn arrives with the bridge token, and the server runs the emergency rules on what the person said before anything else. An emergency or crisis interrupts with the 911 or 988 message and is written to the record.
- Answer through the gateway. The server asks the model through the gateway, with the guide's rules and the record brief, under that person's own consent. The gateway strips identifiers before any text leaves and logs the call. No vendor key sits on the service, which refuses to start if one is set.
- Filter before it is spoken. The answer filter reviews the reply and replaces any sentence that crosses the line before the agent speaks it.
- Fall back without losing the guardrails. With no microphone, the app falls back to the browser's own speech and the typed guide, which take the identical path through triage, the gateway and the filter.
MY TALISMAN
Ask, with graded evidence
A patient can ask anything medical and gets information with a reference behind every statement, a grade on every source, a warning when a source is weak, and a score from 0 to 100 capped when the research is thin. It is never an opinion about the patient and never a diagnosis.
Any question, including the ones nobody else will take
Patients ask about conditions, symptoms, risks, tests, drugs, supplements, herbs, diets, acupuncture, devices and treatments that are not yet standard. Ask takes all of it. Refusing the question is the usual approach and it fails: a patient told to consult their physician goes and asks somewhere that answers, and the place that answers is almost never the place that grades its evidence.
What every answer carries
- A reference on every statement. Any sentence without a source is dropped before the patient sees it.
- A grade on every source, with the reason for the grade beside it.
- A plain warning when something is weak, saying how many are weak and how weak: "one of these is a news interview, not a study, so treat it as a lead rather than evidence."
- An evidence score from 0 to 100, built in the style of the NERD score, with hard caps.
- Facts from the patient's own record, restated without interpretation.
- Up to three questions back, such as whether the patient feels cold more easily than people around them. Those answers go into the record and are never used to reach a conclusion about the patient.
- Questions for the doctor, what research would settle it, and studies recruiting now.
References are written, not spoken
Every reference links to its source and is also written into the chat as a clickable search, so the patient can run the same search and see the source in its own context rather than taking the app's word for it. A citation a patient cannot check is a decoration. In voice mode the guide never reads references aloud; it says once that they are in the chat and each is a link, because a spoken list of journal names, volumes and authors is unusable and crowds out the answer.
Source grades
Every source gets one grade, with the reason beside it, depending on the kind of study and what it found.
| Grade | What the patient sees |
|---|---|
| Solid | Strong support for the statement |
| Fairly solid | Support, with limits |
| Thin | Little research behind it |
| Sketchy | Called out in the warning line |
| Very sketchy | Called out in the warning line |
| Unverified | The source could not be checked |
| Retracted | A withdrawn paper, named so nobody relies on it |
News reports, interviews, personal accounts, health media, animal studies, preprints and retracted papers are always called out by type. News stories and interviews are welcome and graded as what they are, because a patient who heard something on the radio needs that claim engaged rather than ignored. Ask also reads conclusions in the correct direction: tests hold it to reading "did not significantly reduce" as a null result and "the evidence is incomplete" as unsure. A system that cites a paper for the opposite of what it found is worse than one that cites nothing.
The score, and the caps that keep it honest
The score is built from the kinds of studies available, the amount of focused research, agreement between conclusions, recency, and how directly the research answers the question asked. Parts that cannot be assessed are labeled not assessed rather than given a middle value, which is how most scoring schemes quietly inflate. Funding and conflicts of interest are not yet scored, and every answer says so.
Two caps sit above the score and cannot be argued past. A question too broad to match specific research never scores above 35, so "is turmeric good for you" does not come back looking settled. A question with no focused systematic review and fewer than two focused randomized trials never scores above 50, so a score above 50 means a review or two trials exist on the exact question asked, a claim a reader can check. The caps exist because a score built on the volume of research rather than its quality is the standard way an evidence score goes wrong. A test holds them on every release, and every answer lists the research that would raise its own confidence.
One set of rules for everything graded
Conventional treatments, alternative treatments and products from affiliated companies are graded by the same rules, affiliations are disclosed, and no referral payments are taken. Each part is load-bearing in a different direction.
Against conventional medicine it means the weakness of standard practice gets stated too. Publication bias, retractions and negative results that were never published are real, and a system applying skepticism only to the challenger is not a scoring system. A standard treatment supported by two small industry-funded trials scores like one.
Against alternative medicine it means a mechanism is engaged on the evidence rather than dismissed for lacking consensus, then graded without mercy. A supplement with three underpowered trials gets Thin and a capped score.
For the company's own products it is commercial as much as intellectual: a product graded under rules its maker does not control is worth more to a regulator than evidence that maker scored itself.
Where it stops
- "Do I have X?" returns what the research says about X's signs and screening, the fact that only a clinician can answer it, and a question to bring to that clinician. Never a yes or a no.
- No ranked list of likely conditions, and no treatment selected for the patient. The boundary sits where the federal enforcement-discretion examples for symptom checklists and herb-drug interaction lookups sit.
- Crisis wording routes to 988 and emergency wording to 911 first, through rules a model layer may escalate and can never soften. Whether fish oil helps after a stroke five years ago is recognized as research rather than an emergency.
Only the topic leaves
Outside searches carry topic terms from a medical vocabulary and nothing else: conditions, symptoms, treatments, supplements, drug names, laboratory names from the record's registry, and terms from the person's own record. Leftover words are never searched, and a term sharing a word with the person's or a family member's name, postal code or birth date is dropped.
Searches reach six allowlisted hosts: the European and United States biomedical literature, the national consumer health library, the federal trials registry, the federal drug label database and a news search. Every request passes through one egress module that refuses all of them unless lookups are switched on, refuses any other host or anything unencrypted, and sends through a proxy. With lookups off, Ask answers from the cached library and lists the outside sources as unavailable.
What the discussion adds to the corpus
Every question asked and answered deposits something hard to buy. The library accumulates graded findings by topic, so the second patient to ask gets a faster and better-sourced answer than the first, and the corpus grows along the topics patients care about rather than what a content team guessed at. Where the research is thin, the gap is recorded with the study that would close it: a map of what medicine does not know, assembled from what patients keep asking.
The patient's own questions are written to their record, and the morning brief includes what they have been asking about, often the first signal that something is wrong before any test is ordered. Their follow-up answers and reports of what happened after a treatment enter as their own observations, under their signed research authorization before any research use. Those are independent judgments that did not come from the system's own answer, the input a record learning how far each source can be trusted cannot generate for itself. A feature most products treat as a cost center is here the supply line for the asset.
MY TALISMAN
The whole family, and what happens at a visit
A parent carries their children, and any patient carries the people they hold legal medical responsibility for, inside one account. At the visit, the guide announces that it takes notes without recording, then hands the clinician an editable synopsis while everything else is discarded.
One account, separate records
The patient switches between the people they care for in one tap. Each has their own record, and nothing is merged into the patient's. The account holds a proxy link to each one storing the legal basis, the signed authorization and the end date.
Keeping them separate is the difference between a family feature and a liability. A merged record cannot be unwound when a child turns eighteen, cannot withhold the care a minor may lawfully keep private, and cannot be divided when a custody order changes.
| Person | Legal basis | How access starts | When it ends |
|---|---|---|---|
| Minor child | Parent or legal guardian | The parent signs the proxy authorization | At 18, when the guide hands the record to the young adult |
| Spouse or other adult | That adult's own grant | The adult signs, at full record, view only or chosen sections | When the adult revokes it |
| Adult who cannot decide for themselves | Power of attorney or guardianship | The patient uploads it; a reviewer confirms it before access opens | When the document ends or is revoked |
An adult cannot be added as a child, and a test holds that. Parent access ends on the child's eighteenth birthday, leap-day birthdays included.
Legal responsibility, proven rather than assumed
A proxy grant is never implied by a role. Each is a record carrying its legal basis, its signed authorization and, for a minor, its end date, sealed at rest. Access moves only after proof: a one-time code to a contact already on file rather than one the requester typed, staff identity proofing where no usable contact exists, the adult's own approval from their own account, and a birth certificate or court order recorded by staff for a minor. Neither the requester nor the staff can approve access to an adult on that adult's behalf. A grant made from typed details never opens an existing record: the person added gets a record holding only what was typed, so nobody can use that screen to learn who is in the system.
- Care a minor may consent to alone. Records state law lets a minor keep from a parent are withheld from the parent's view, and the consent engine enforces it on every read rather than at the screen.
- Shared custody. Another parent or guardian may request their own access, and a parent whose rights are limited by a court order uploads the order.
- The guide speaks about that person by name while a proxy manages their record, and never mixes their information with the account holder's.
- Every proxy action is logged in the dependent's own access log, which an adult dependent can read.
- Consents are per person, made on that person's own screens.
One account carrying a spouse, two children and an aging parent is four lifetime records and four streams of outcome data from one enrollment. Adult children managing a parent's care are the most motivated users in health software and the worst served.
The visit notetaker
The night before, the guide delivers a prep brief: why the patient is going, what changed since the last visit, and the questions they raised in Ask. During the visit it takes notes and keeps no recording. The clinician edits or accepts them; everything else is destroyed.
- It announces itself. When the visit starts the guide tells the clinician, in its own voice, who it is, whose guide it is, that it would like to take notes so the patient can remember the discussion, that it will not record audio, and that it will show the notes at the end to edit or accept. In typing-only mode the clinician reads the same words on screen.
- The clinician consents or declines. A tap or a spoken yes, logged with their stated name, role and the time. A clinician who declines gets a guide that stays silent for the whole visit.
- It takes notes. Speech becomes text only long enough to write them, and a banner reading "taking notes, no recording" with the consent time stays at the top of the screen.
- It hands over a synopsis: what was found and recommended, medication changes, follow-ups and the questions answered, as an editable document.
- The clinician accepts, edits or discards. Only the accepted synopsis enters the record, marked as approved by that clinician and labeled a summary for the patient's own use rather than clinical documentation. Audio, transcript and drafts are deleted, and tests confirm a discard keeps nothing.
Announcing it out loud is the design rather than a courtesy. Ambient scribe products are consented by the clinic that bought them; this one is in the patient's hand, pointed at a clinician who has agreed to nothing. The announcement and the yes serve as consent under all-party consent laws, which can treat live transcription as recording even when nothing is kept.
Keeping no audio removes the asset that makes clinicians refuse. There is no recording to be subpoenaed, leaked, replayed or compared against the clinician's own note, and the only durable artifact is a document they read and approved. At a clinic running Arca, the clinic's own notetaker takes the visit and the patient's guide steps aside. Afterward the accepted notes become a plain-language summary with tasks and timed check-ins.
Telemedicine in one tap
The patient taps see a doctor now, or the guide offers it when a symptom calls for one. The guide asks two or three questions, builds a visit packet from the record, and stops. The packet holds the reason for the visit, medications and allergies, ninety days of laboratory results and relevant history, with the full record included only if the patient chooses it. The patient approves it and signs the clinician's visit consent, with any fee shown first. Afterward the clinician's note, any prescription and the follow-up plan come back into the record as reminders and check-ins.
The patient's own practice is offered first whenever it offers telehealth, then the partner's own clinicians, then a contracted national network supplying licensed clinicians, state licensing, malpractice coverage and prescribing under the partner's brand. Offering the patient's own doctor first is a retention decision: no practice will hand its patients a product that routes them away. A company-owned network can follow once volume supports it, structured as corporate practice of medicine rules require. The guide prepares and hands off; the licensed clinician diagnoses and prescribes.
The agreements, walked through
Six agreements set the limits of the service and give every intended use of the data a signed basis: the terms of use, the privacy notice, the research authorization, the notetaker notice, the consumer health data consent with separate signatures for collection and sharing, and the proxy authorization. The guide walks the patient through each and never accepts on their behalf.
The research authorization is deliberately separate, with its own screen and its own signature. Health data consent bundled into general terms is the kind regulators and courts most often refuse to enforce, and a separate signature is what lets each intended research use stand. The patient chooses care only, care and research with identity held apart, or care and research with outside projects that receive results rather than records. Care never depends on signing it.
Every signature, from a patient or a clinician at a visit, is stored as a permanent tamper-evident record holding the signer, the document version, a fingerprint of the exact text shown, every choice made, the device and the time. The chain's head is anchored outside the database, so a chain rewritten inside it still fails against the anchors. One step produces the signed documents and the proof for an audit years later.
MY TALISMAN
The screens
Seventeen designed screens, sixteen for the phone and one for the desktop, grouped into the flows a patient runs: getting started, every day, the visit, and the wider screen. Every one follows a single invented patient and is marked as sample data.
The screens use the direct edition's own palette and typefaces, with a partner's colors swapped in by its brand pack. The guide's avatar is two concentric circles in the brand color, the same mark wherever it speaks. Evidence chips appear wherever the guide makes a claim, so the strength of a statement is visible in the same glance.
Getting started
Five steps, each labeled so the patient knows how much is left, run as a walkthrough rather than a form to fill in alone.

| Screen | What it shows | Main action |
|---|---|---|
| Welcome and voice | The partner's logo with "with My Talisman" beside it, the guide's greeting, two voices to choose from, the talk-or-type choice, and a link for existing accounts | Choose a voice and begin |
| The guide names itself | "From what you've told me, I'd like to be called Wren," an editable name field, five alternatives, and a way to say a name out loud instead of typing it | Accept the name or give another |
| Terms walkthrough | Four short agreements with their status, the short version of each in three bullets, an invitation to ask the guide anything before signing, and a sign control | Sign each agreement in turn |
| Upload walkthrough | The guide offering to add the patient's last laboratory report together, numbered steps, camera guidance that says when the page is in frame, and four ways in: take a photo, choose a file, connect my clinic, say it | Add the first document |
Putting the agreements inside the walkthrough rather than behind a scroll-and-accept box is what makes the signature worth having. The patient asks what a clause means before signing, and the questions they asked are stored beside the signature.
Every day
The everyday screens are built around a claim the app earns each time it opens: something was noticed since the patient last looked.



| Screen | What it shows | Main action |
|---|---|---|
| Home | Three things the guide noticed, each with an evidence chip, today's medications and visit, the prep brief, see a doctor now, and an avatar for each family member | Open a finding |
| Ask | The answer with a reference on every statement, a grade and a reason on every source, the warning line, the capped score, the facts from the patient's own record that bear on the question, and up to three questions back | Read a source, or answer the guide |
| Record timeline | Every entry by date with its source, filters by kind of care, the original document behind each value, and care outside conventional medicine alongside the rest | Open an entry and see the original |
| Medications | An interaction card with its evidence grade and a see-why control, every medication and supplement with its timing and refill countdown, add by voice, scan a bottle | Mark a dose, or add a medication |
| Result explained | One value against that laboratory's own range, charted across its history with the date a treatment started marked, a plain explanation, add this to Thursday's questions, and the original report | Add the question to the next visit |
| Scan viewer | A full imaging series with scroll, contrast, zoom, measure and compare, the radiologist's report rewritten in plain words beside it, and send to a doctor | Scroll the series, or send it |
| Photo tracking | A body map with each tracked spot, a monthly series with a measured width on each photograph, the standard warning signs, and send to a dermatologist | Take this month's photo |
| Show my doctor | The same series full screen for a clinician, with scale, a longest-width line across the months, a month slider, and the line "sizes are approximate, no analysis applied" | Hand the phone to the clinician |
The photo screens are where the regulatory line is visible in the interface. The app measures and presents, says the sizes are approximate, and states that no analysis was applied. It never offers an opinion about what the spot is, because that is a diagnosis and belongs to a licensed clinician.
The visit
Three screens carry the visit: the brief before it, the notetaker during it, and the route to a clinician when there is no appointment to wait for.

| Screen | What it shows | Main action |
|---|---|---|
| Prep brief | Why the patient is going, what changed since the last visit, the questions they raised in Ask, and the questions the guide suggests, each addable by voice or typing | Add a question of their own |
| Visit notetaker | The no-recording banner with the consent time, suggested questions as the conversation moves, the notes so far, and a control to end the visit and show the clinician the notes | Hand the synopsis to the clinician |
| See a doctor now | What is going on, exactly what the clinician will see, the next clinician with the wait and fee, share and start video, call my own doctor's office instead, and the line to call 911 | Approve the packet and start |
| New diagnosis research | A plain overview, treatment options with their evidence strength and a control explaining the scoring, recruiting trials with site, distance and a call packet, and the line "information, not a diagnosis, decide with your doctor" | Take the call packet for a trial |
| Share with a doctor | The recipient, what to include, a link that expires in seven days, everyone holding access with a control to remove each, and a log of everyone who looked | Send the link, or revoke one |
Showing the patient exactly what a telehealth clinician will see, before they approve it, is the part most products skip.

The wider screen
The desktop web screen puts navigation down the left side, the day in the middle with the family switcher above it, and the guide's conversation in its own column on the right, where it never covers what it is talking about. The sample exchange shows the guide explaining that iron taken with morning coffee absorbs less and that newer research supports every-other-day dosing, then offering to add both to Thursday's questions.
The desktop screen is not a courtesy to older patients. It is the screen a caregiver managing three records uses, the one a patient reads an imaging series on, and the one a front desk uses to finish an enrollment. Sign-up happens in a browser, so it is the first screen most patients ever see.
MY TALISMAN
The My Talisman market
Consumer AI health assistants became a mass market in 2026. My Talisman competes as the branded, HIPAA-grade guide an organization hands its own patients, over a lifetime record that keeps every fact with its source, with family records, imaging and graded evidence behind it.
The market My Talisman sells into
Patient engagement and personal health records, worldwide.
The demand is settled, the access is not
Whether patients want this is settled by public figures. ChatGPT Health opened to all United States adults in July 2026 and handles about 300 million health questions a week. Microsoft reports more than 50 million a day, each asked of a system that knows nothing about the person.
Against those 700 companies, fewer than 50 provider organizations had pledged. Clinics are the bottleneck rather than patients, and that gap is the distribution thesis: reach patients through the organizations that treat them, on the side with less competition. The federal Medicare application library has opened a category for conversational AI assistants and limited it to applications meeting HIPAA, verified identity, the industry code of conduct and connectivity to the aligned network. Five were listed in June 2026. The standard excludes assistants outside HIPAA, which is most of the ones holding those question volumes.
Sizing, and what the published numbers are worth
Published estimates for patient engagement and personal health records run $33.8 billion to $41.2 billion worldwide and about $14.6 billion in the United States. Estimates for the narrower personal health record category run from under $50 million to over $6 billion depending on the definition, a hundredfold spread and a reason to rest no plan on it.
Plus 6,436 ambulatory surgery centers and more than 2,700 direct primary care practices, counted by Definitive Healthcare, the federal payment advisory commission's March 2026 report and DPC Frontier. Each brings its own patient panel on the day it signs, so the unit that matters is a panel rather than a download.
Who pays
Patients never pay to own their data, and records are never sold, rented or transferred. Organizations pay.
| Buyer | What they receive | Structure |
|---|---|---|
| Clinic groups, employers, health systems, wellness companies | A branded edition from one code base, set by a brand pack: logo, colors, welcome message, care team, feature switches, and usage reporting with nobody identified | Per active patient per month |
| Arca clinics | My Talisman as the clinic's own patient portal, at no cost to its patients | Included in the Arca subscription, per clinician per month |
| Outside application developers | Consented access through the standard interfaces, reading only what their application created | Per connection |
| Individuals | The direct edition under the My Talisman name and colors | Undecided; the terms require a shown price and agreement first |
Prices are set with the first partners during the pilot. Including the application with Arca rather than charging a clinic's patients is a deliberate trade: the portal is what a clinic already pays a vendor for, so giving it away converts an acquisition cost into a feature of a subscription already sold.
One code base produces every branded edition, and the white-label route is the opening rather than a side business because of a constraint on the buyers. A clinic group, an employer or a health system cannot hand its patients a chat product outside HIPAA, and will not hand them one carrying a technology company's brand instead of its own. The organizations with the patients cannot use the products with the capability. A partner wanting an application built to its own specification gets a separate one, because every improvement to My Talisman reaches all editions and a fork would end that.
Where it wins
| Competitor | Overlap | What it does not have | Scale |
|---|---|---|---|
| ChatGPT Health | Records and laboratory results, results explained, visit prep | Records through an aggregator; outside HIPAA; no imaging, family records or white label; sued July 2026 over advice not to see a doctor | All US adults; 300M a week |
| Microsoft Copilot Health | Records from 50,000-plus organizations, 50-plus wearables | Requires Microsoft 365; records through an aggregator; no HIPAA claim | Preview since May 2026 |
| Google Health Premium | Records, wearables, an AI coach, QR sharing | Built around one wearable brand; no white label, no imaging | $9.99 a month |
| Apple Health | Health records on the phone, AI insights | One phone platform; on the device; no write-back to a clinic | Redesign announced September 2026 |
| Amazon Health AI | An agentic assistant, HIPAA-compliant | A funnel into its own clinic and pharmacy | All US customers since March 2026 |
| Counsel Studio | White-label AI care for plans and apps | Physician-supervised and diagnoses; no lifetime record or imaging | $36M raised |
| Verily Me | A consumer record, an AI agent, a research registry | No clinic software or imaging platform | $300M raised, March 2026 |
| Epic MyChart and its agent | A portal with proxy access and a text agent | Stops at that vendor's own customers | 115M patients |
| Guava Health | Portal sync, family profiles, visit prep | No network retrieval, imaging or white label found | $78 a year |
| Citizen Health | An AI advocate over a record with imaging | Rare disease only | $44M raised |
Six advantages hold across that table. An organization's own brand over a HIPAA-grade service. A record kept continuously across every doctor and every kind of care rather than fetched one question at a time. Graded evidence, which no competitor reviewed publishes. Family records, imaging and a notetaker that keeps no audio, in one application. Distribution through care, reaching a clinic's whole panel at once. And neutrality, since every option is scored by the same rules.
The platforms have the patients' questions. The clinics have the patients.
What the patient relationship is worth
The application earns a per-patient fee, and that is the smaller half of its value. The record under it learns how far each source can be trusted, from independent judgments about what happened. Claims data arrives late and describes billing. A clinic note describes what the clinician concluded, the system's own answer coming back to it. A patient answering a check-in after a new treatment reports something the system did not produce. A record surrounded by applications has a supply of those judgments; a record sold alone does not, and its weights sit near their starting figures indefinitely.
Two further consequences follow. A patient holding their own lifetime record does not lose it when they change doctors, so the record outlives every clinic contract in it and the patient becomes the reason the next clinic connects. And a patient enrolled through one organization carries their family, several lifetime records from one enrollment.
The relationship also opens an adjacent product on the same core: a pregnancy application that reaches the most motivated health users there are and produces a second lifetime record at birth. Scoped, not built.
Risks we raise ourselves
| Risk | The answer |
|---|---|
| The platforms reach more people in a day than this will in years | Reach patients through the organizations that treat them, each bringing a whole panel. The platforms cannot carry a clinic's brand or meet the HIPAA standard the Medicare library requires |
| AI health tools draw federal and state scrutiny | The wellness line, triage before every answer, the filter after it, and every feature classified against the January 2026 federal guidance before it is built |
| Name conflict | A knockout search found a medical billing company using the name. Counsel clears it before public use, and it stays on the patient-facing app only |
| A consumer AI tool says something it should not | Every answer passes triage and the filter, whose limits are measured rather than asserted. Vendor agreements with zero retention come before any real record |
The sequence has three gates. Now: a hardened web application and the clinic portal on invented data. Before the first real patient: the security gates, signed vendor agreements for language models and voice, and a pilot with the first partner. Then: native phone applications, certification under the industry code of conduct, and a Medicare library listing, which requires verified patient identity through a contracted proofing partner.
ARCA
Arca, the clinic system you talk to
Arca is the medical record and practice management system a clinic runs by speaking to it. A guide listens, drafts, looks things up, books, codes, bills and follows up, and a person approves anything that changes a chart, sends an order or submits a claim. One code base serves every outpatient specialty, from direct primary care to imaging and surgery centers, in a human edition and a veterinary edition.
What the product is
Everyone in the clinic has a guide, on a workstation, a tablet or a phone, by voice or by typing. Behind the guide sits a team of agents, each with one job, its own permissions and its own log: a receptionist that answers the phone, a registrar, a chart summarizer, a notetaker, an order stager, a safety checker, an inbox agent, a coder, a claims agent, a prior authorization agent and an analyst. The guide does the work before anyone asks for it. Charts are summarized before rooming, eligibility is checked the night before, orders are staged out of the conversation in the room, the claim is coded at signing, and inbox replies are drafted before anyone opens the inbox. One code base serves every clinic type: a clinic picks a skin for its specialty, which sets its templates, flowsheets, order sets, billing rules and widgets. Finished clinical facts go to the lifetime record, working state stays in the clinic's own store, and each clinic's patient accounts sit in its own ledger under its own key with no path to another clinic and no path to research.
Screens remain for anyone who prefers them, and every screen can be driven by voice, by touch or by keyboard. Arca runs in any modern browser, so one code base serves Windows machines, touch-screen workstations, Macs, iPads and Android tablets, and installs as an application on each. On a tablet a clinician opens on a rounds screen with the patient list beside the chart, and patient data is never stored on the device.
AI drafts, retrieves, checks and flags. Licensed people diagnose, prescribe and sign.
Every agent action waits for a person's approval
No agent signs a note, places an order, sends a prescription or submits a claim without a person approving it in the approval queue. The guide shows what it heard and what it will do, and the person accepts, edits or refuses. A clinic may switch on automatic handling for one named low-risk action, such as an appointment reminder or posting a clean remittance, and that switch is recorded. Every action carries the approval that allowed it on a tamper-evident log. The rule also closes the most obvious way to attack a system run by conversation: outside documents, faxes and patient messages are treated as data, so text inside them cannot instruct an agent.
Two design decisions that are commercial arguments
Physicians rate their electronic health records at 45 out of 100 on the standard usability scale, in the largest published study of the question (Melnick and colleagues, Mayo Clinic Proceedings, 2019). The average across all kinds of software is 68, and 84.1 puts a product in the top four percent of everything measured. Forty-five is a failing grade for the software a physician spends more of the day in than with patients.
Arca's target is a mean of 90 or higher for every role: clinician, nurse and medical assistant, front desk, biller, manager and patient. No release ships if any role's mean falls below 80, and any single score under 70 opens a defect. The survey is the ten standard items, scored the standard way, run quarterly on a random sample at each clinic.
Nobody sits through training. Every function in Arca is registered in one catalog with its roles, its steps, the screen it lives on and the phrases that invoke it, and a release fails if any route lacks an entry. From that catalog the guide will tell a user the steps, show them on the screen with each element outlined in turn, do the function and bring it back as a pending approval, or let them practice in the clinic's own practice copy, built from the real staff list so everyone practices as themselves. A function does not ship unless a test proves the guide can teach it, and help pages are generated from the same catalog so the guide and the written help cannot disagree. The target is a front desk hire working on real patients in the first hour and a clinician finishing a first real visit with the guide in the first session, measured as time to independent work by role.
The second decision is about interruption. Drug interaction alerts are overridden about 90 percent of the time in published studies, which means the alert has stopped carrying information and become a tax on attention. In Arca each clinician chooses, for each kind of alert, whether it interrupts, sits in a side panel, arrives in a daily digest or stays off. A short list of true safety alerts stays locked on because the clinic carries the liability: severe allergy, a contraindicated pair, a dose above the maximum, a pregnancy contraindication, a critical laboratory value, duplicate high-risk therapy. Those are overridden only with a reason, written to the log, and are never demoted. Everything else is tuned. Any clinician-controlled alert accepted less than 30 percent of the time over 90 days, with at least ten decided interruptions, stops interrupting on its own and moves to the side panel, and the clinic's alert owner is told and can restore it. The budget is five interruptive alerts per clinician per day.
Both decisions are commercial because both are measured per release. A buyer comparing systems has no usability number from an incumbent and no alert acceptance rate.
Who buys first, and in what order
The entry market is the one with the lightest rules and the thinnest competition. Cash-pay primary care, direct primary care, concierge, wellness and integrative clinics do not bill Medicare, so they need no federal certification to buy, and the systems they use today offer a bolted-on scribe at $290 to $770 a month. More than 2,700 direct primary care practices operate in the United States, counted by DPC Frontier. Small practices on the lowest-rated incumbents are the second group, sitting on unstable vendors and reporting the worst satisfaction in the field.
| Order | Buyer | What opens it |
|---|---|---|
| First | Cash-pay, direct primary care, concierge, wellness and integrative clinics | A card processor and the clinic security program. No federal certification required |
| Second | Insurance-billing primary care, internal medicine, chiropractic, podiatry and urgent care | A clearinghouse contract, payer enrollment, the federal edit files and a coding license |
| Third | The 395,000 US group practices, 11,000 urgent care centers and health-system-owned clinics | Federal certification, record-network participation and outside ratings |
| Fourth | 15,000 imaging centers and 6,436 ambulatory surgery centers | Facility claims at volume and each scanner vendor's worklist check |
Counts from Definitive Healthcare and, for surgery centers, the federal payment advisory commission's March 2026 report. Multispecialty groups and management services organizations are where the single code base is worth the most, because they run several specialties on several vendors today and no incumbent covers primary care, pediatrics, orthopedics, women's health, chiropractic, wellness, concierge, urgent care and imaging on one patient record. Hospitals are not the first target, and the system is built to run them.
Outside the United States the order is set by how much regulatory work stands between the software and a paying clinic. English-speaking markets with private cash-pay clinics come first: the United Kingdom's private sector, Australia and Canada. Private hospital groups in the Gulf states, India and Latin America follow, where a full incumbent installation is out of reach on price, and where 882 hospitals operate in the Gulf alone by the GCC Statistical Centre's count. The European Union follows once its new conformity process takes effect from March 2027. National public-sector systems come last, because their incumbents hold decades-long positions and their procurement runs on a different clock.
The clinic system, on screen






Screens from the clinic system running on invented clinics and invented patients.
ARCA
Charting, orders and the room
A notetaker that drafts the visit from the conversation and keeps no recording, notes whose every sentence links to the moment it came from, prescribing with controlled substance signing, laboratory orders and results on the real standards, alerts each clinician tunes, and a clean line between the finished facts that go to the lifetime record and the working state that stays with the clinic.
The in-room notetaker
The notetaker is the default way to chart. The guide announces itself to everyone in the room, says it takes notes and keeps no recording, and records each person's consent with the time. If anyone declines, the clinician charts by voice after the patient leaves, or by typing. Speech becomes text only to build the draft, and the clinician can speak to the guide mid-visit to stage an order without breaking the conversation.
Within seconds of the visit ending the note appears in the clinic's template, with staged orders and prescriptions, suggested diagnoses and codes marked as suggestions, patient instructions and follow-ups. The safety checker flags anything the note states that conflicts with the chart, an allergy, a medication, a side of the body, and anything said that the note left out. After signing, the working transcript is deleted.
Keeping no recording is a commercial position, not a technical shortcut. Patients agree to ambient notes at high rates when told the basics and at far lower rates once they hear about AI and data storage: 81.6 percent against 55.3 percent in a 2025 study in JAMA Network Open, a fall of 26.3 points traceable to storage alone. Scribe vendors differ, some keeping audio 30 to 90 days and others none. Arca keeps none, which removes the concern rather than disclosing it. Nine states require every party's consent before a conversation is transcribed, and the announced consent covers them.
Every sentence linked to its moment
Each sentence the draft puts in the note is linked to the spans of the visit that support it: the segment, the character offsets, the time, the speaker, and the words as spoken when the visit was translated. A clinician touches a sentence and reads that moment with the supporting words highlighted. A sentence with no supporting span is marked for review, and a number or dose that nobody said is flagged on its own. The linking is deterministic word matching that runs offline, finding lexical support only, so a sentence supported in meaning but worded differently is marked for review, which is the safe direction to fail in.
The published error studies are why this exists rather than a marketing feature. A randomized trial of 238 physicians across about 72,000 visits found an ambient scribe cut note-writing time 9.5 percent, roughly 41 seconds a note, with occasional clinically significant errors, mostly omissions (NEJM AI, December 2025). A separate study found invented content in 1.47 percent of sentences and omissions of 3.45 percent of what was said (npj Digital Medicine, 2025). Controlled studies put the saving at one to two minutes a visit, far below vendor claims of 45 to 72 percent. At signing Arca computes how much of the draft the clinician changed and writes it to the log, which makes the quality of the drafting measurable rather than asserted.
Arca also proposes structured chart entries from the transcript: problems with their diagnosis codes, medications with ingredient, strength and form, allergies with the reaction, vitals, orders and the follow-up. Each carries the moments it came from and an action against the chart: add, change, stop, or already present and shown but never written. Two extractors run, a rules extractor always on and offline and a model extractor only under live consent, whose entries are merged under the rule entries and never over them. Visits run in 17 configured languages.
Content packs by specialty, and the reviewer rule
A specialty's clinical content ships as one versioned document: note styles and templates, structured exam elements, flowsheets, order sets, rating scales, device import mappings, coding rules and patient instructions. Five exist: primary care, concierge and direct primary care, wellness and integrative, urgent care, and orthopedics. A pack counts as reviewed only when it names a practicing specialist with credentials who reviewed that exact version, and a coding owner. A live clinic cannot turn on an unreviewed pack, and the refusal is recorded. The five shipped packs name no reviewer and show as unreviewed sample content wherever they are used. A clinic's changes sit in a layer above the pack, so an upgrade changes the pack underneath without discarding what the clinic set.
Prescribing and controlled substances
A prescription is drafted from drug, strength, dose, frequency, days supply, refills, substitution, diagnosis and pharmacy, and Arca writes the directions and the quantity. A real-time benefit check shows the patient's 30-day cost under their own coverage, with up to three alternatives in the same class, before signing. Prescriptions are written in NCPDP SCRIPT 2023011, the only version allowed from January 1, 2028.
Controlled substances follow the federal rules in full. The prescriber is identity-proofed to the higher federal assurance level. A manager or owner who is not the prescriber turns the capability on, which is the two-person control the rule requires. A check of the state monitoring program within 24 hours is required before signing. Signing takes two different kinds of factor: a session minted from a hardware security key within the last 120 seconds, for the same person and clinic, plus the prescriber's signing code. A phone code is not offered, because a phone is a second thing you have rather than a second kind of factor, and a face or fingerprint passkey signs a clinician into Arca without being able to sign a controlled prescription.
Every drug and result check fires through the alert tiers. A locked safety tier always interrupts: severe allergy, a contraindicated pair, a dose above the maximum for age, weight or kidney function, a pregnancy contraindication, a critical laboratory value, duplicate high-risk therapy. Those are overridden only with a reason, written to the log, and are never demoted. Everything else each clinician sets to interrupt, side panel, daily digest or off, down to a floor the clinic sets. Every alert says why it fired, links to the drug entry and to the record event behind the patient fact, and offers one fix.
Laboratory orders and results
Orders carry coded tests and diagnoses, taken from the problem list when none are given, and run over two-way HL7 v2.5.1 interfaces with acknowledgment. Medicare medical necessity flags decide when an advance beneficiary notice is needed, and the patient's choice is recorded before the order goes out. Results are matched to the order and to the patient on record number, family name and birth date, and a mismatch is held in the inbox and rejected rather than filed. A corrected value links to the one it replaces, which is marked superseded.
Release to the patient follows the federal information blocking rules rather than clinic habit. Each result releases at its arrival plus the clinic's hold setting, which starts at zero hours and cannot exceed 72. A clinician can add a note the patient will see, release early, or hold one order's results under the preventing-harm exception with a logged reason. Blanket delays are presumptively unlawful and the product does not offer them.
Where the facts go
Finished clinical facts live once, in the lifetime record, and everything else reads a view of them. The clinic is the legal custodian of its chart and must keep it, answer subpoenas and access requests, and take all of it when it leaves. When a patient uses their own application, a copy of each finished record flows into the patient's own domain under their federal right to direct copies to an application of their choice, and each domain has its own keys, so erasing the patient's copy never touches the clinic's chart. Sensitive entries carry labels the consent engine enforces on every read: substance use records, care a minor may consent to alone, reproductive and gender-affirming care where state law requires segregation, and psychotherapy notes.
ARCA
The front office
Scheduling, registration, forms, the phone line, the inbox, telehealth and the patient portal. The guide answers the phone, books, texts, fills cancellations, sorts the inbox and drafts every reply, and a person approves. The portal is not a separate website: it appears inside the patient's own application under the clinic's brand.
Scheduling
The schedule books people, rooms and equipment together, because an outpatient appointment is rarely one resource. Clinicians, rooms, chairs, procedure rooms, imaging modalities, audiology booths and operating rooms each carry their own hours. One appointment books all of them at once, all or nothing, and is refused with the list of conflicts when any one is already taken. The check and the write run under one lock per clinic, so two bookings cannot both pass.
Visit types carry length, resources, prep instructions, forms, eligibility checks and the templates they open. Staff book by saying it: put her in her surgeon's first opening next week, 30 minutes, knee injection. The guide finds the slot, checks the rules and books after approval. Recalls come from care gaps, overdue follow-ups, specialty cycles and membership benefits, and a cancellation texts the first matching person on the waitlist, filtered by clinician, visit length, date window and consent.
Registration and the forms
Intake and consent forms go out by link, are completed by voice or on screen, and are signed in the patient's application, on the web or on an office tablet. The guide reads and explains each form and never signs for the patient. A signature is stored with a cryptographic fingerprint of the exact text signed, so what the patient agreed to is provable rather than inferred from a version number. At arrival a rotating care code in the patient's own application identifies them, and insurance card and identification photographs are read and compared with the file.
Eligibility and benefits are checked in batch the night before and in real time at booking, and the answer shows beside each appointment: active or not, plan, copay, deductible and what is left. Members and self-pay patients are marked as needing no insurance check. Prior authorization is checked at booking, and the balance and copay are presented for payment before arrival.
The phone line
The guide answers the clinic's phone and text lines, and the order of operations matters more than the capability. Every caller sentence passes the emergency rules before anything else reads it. An emergency or a crisis ends the automated call at once: the caller hears the 911 or 988 instruction, the call passes to a person with a warm transfer, a safety item goes to the inbox and the escalation is logged.
Nothing about a patient is said until the caller gives a date of birth and then a six-digit code texted to the mobile number on file, never to the calling number unless that is the number on file. The guide answers identically whether or not the details match anybody, so a caller cannot learn who is a patient by probing, and three failed attempts pass the call to a person. Inside those walls the guide handles hours, billing questions, booking, rescheduling, refill requests and balances, in English or Spanish.
Texting runs on recorded consent per patient with its source and time, and reminders, recall, reactivation and waitlist offers go only with consent. STOP opts the number out at once, and texts carry the clinic's name and a time and nothing clinical. Other modules queue patient messages in one outbox and never send anything themselves, so a message without consent, after a STOP or without a mobile number is held with the reason rather than lost.
One inbox
Results, refill requests, portal messages, faxes, referrals, texts and calls arrive in one inbox. The guide sorts each item by category, by priority taken from the emergency rules, and by who should handle it, then drafts the reply, the refill or the routing. A person approves, and the role gate holds: clinical items need a clinician, medical assistant or nurse, front desk items the front desk or a manager, billing items a biller. Items per clinician per day and the share the guide drafted are computed from the log.
The inbox is where the clinical argument and the commercial argument meet. In academic primary care from 2019 to 2023, time in the inbox grew 24 percent and time on orders grew 59 percent, and after-hours work did not fall even where scribes were in use (Annals of Family Medicine). A primary care physician spends about 36 minutes in the record for every 30-minute visit (AMA), and a scribe does nothing about either number. In a Stanford pilot, AI-drafted inbox replies cut clinicians' task load and exhaustion scores even though reply time did not change. Arca drafts the action along with the reply.
Public reviews get a drafted reply for a manager or owner to approve, under rules written into the product: a reply never names the reviewer, never repeats anything from the review, never names staff, and never uses words that imply care such as patient, visit or your doctor. These drafts are templates only, because the reviewer gave no consent for a model to read what they wrote.
Telehealth
A clinician starts a video visit from the schedule and the patient gets a text with the link. The patient, or an approved proxy, signs the telehealth consent first and it is written into the chart before the visit starts. The visit is an ordinary Arca visit, so the notetaker runs exactly as it does in the room. Media runs browser to browser with Arca as the signaling server only, over a connection whose credential is a join token good for this visit, this party and this sign-in for at most 30 minutes. Nothing is recorded, and ending the visit revokes every join link.
The patient portal, inside the patient's own application
Arca does not ship a portal website. The clinic's patient-facing functions switch on inside the patient's own lifetime-record application, under the clinic's logo: booking, forms and consents, released results, secure messages, bills, payment plans, membership, refills and video visits. The clinic gets a modern portal it did not build and the patient gets one application instead of five, because the same one also holds their records from every other doctor they have seen.
The two products are deliberately independent. The patient application never depends on Arca, and Arca never requires a patient to install anything: every patient function has a web version, and staff can do anything a patient could do on their behalf. Sign-up needs no application at all. The clinic sends an invitation and it opens in any browser, the patient signs in with a one-time code to the phone or email the clinic already has on file, and the guide walks them through demographics, insurance cards, forms and consents. That removes the step where patient engagement products lose most of their users: a front desk that can send a text message can enroll a patient, and the guide finishes it rather than the staff.
A parent, guardian or caregiver uses their own sign-in and names the patient they are acting for. Access is allowed only with a grant whose holder is that caller, whose identity was proofed to the higher federal assurance level, that is not revoked or expired, and whose scope covers the area being read. That rule is written once and used by the clinical and billing routes as well, and every proxy access and refusal is logged. A proxy request starts a hosted proofing session and never reveals whether the patient details matched anybody.
Arca runs the emergency rules itself on every booking reason, message and form answer, and accepts the application's own screening without relying on it. Two systems reading the same sentence is the point.
The guide answers the phone before the clinic opens, and a person approves everything it did.
ARCA
Getting the clinic paid
A clinic is paid three ways, and Arca handles all three: insurance claims, membership and concierge subscriptions, and cash at the counter. Behind them sits a complete receivables and payables system with every amount held as an exact decimal, the clinic's patient accounts under the clinic's own key, and the clinic's books kept as a separate entity that never sees a patient's name.
The three ways a clinic is paid
| How the clinic is paid | What Arca runs | Who it is for |
|---|---|---|
| Insurance claims | Eligibility at booking, professional and facility claims built from the signed visit, scrubbing, acknowledgments, remittance posting, denial work and appeals | Primary care, specialty, urgent care, imaging and surgery centers |
| Membership and concierge subscriptions | Plans with monthly fees, family pricing, panel caps, employer contracts, recurring billing, member price lists, revenue kept apart from insurance | Direct primary care, concierge, wellness, veterinary wellness plans |
| Cash at the counter | Treatment plans signed before work begins, deposits, checkout and the invoice, payments and the drawer, sales tax, discounts, aging, collections, write-offs and refunds | Self-pay care in every clinic, and the whole of a veterinary practice |
Most systems do one of these well and treat the others as an afterthought, which is why a cash-pay clinic ends up running a claims system with the claims turned off. All three run on the same ledger rows, so an insured patient's share and a self-pay line appear on one invoice.
The claims cycle
Eligibility runs after every booking, from the desk, the phone line, the guide or the portal, and again the day before. A professional claim is built from the signed visit and the coder's approved lines: the note's diagnoses, modifiers and units, the clinic's fees and the rendering clinician's identifier.
The scrubber runs before every send and reports Stop, Warn or Pass with the source it used: required fields and identifier check digits, payer rule packs covering member identifier format, timely filing, prior authorization and coverage policies, modifier logic, procedure pair edits and medically unlikely edits. Remittances post automatically to the clinic's ledger through its own logged write method, idempotent on the trace number. Every rejection and every denied line becomes a worklist item with the dollars at stake and the deadline. The guide drafts a corrected claim, asking a person for what it must not guess, or drafts an appeal letter quoting the signed note's supporting statements. Nothing is resent until a person approves it.
| Measure | Arca target | Industry guidance |
|---|---|---|
| Clean claims on first submission | Above 95% | 90% or better |
| Denial rate | Under 5% | 5% to 10% |
| Days in accounts receivable | Under 40 | Not stated |
In 2025, 41 percent of providers reported that more than 10 percent of their claims were denied, up from 30 percent in 2022. Denial work is where a practice loses money quietly, and it is the work an agent can draft and a biller can approve in a fraction of the time. Coding accuracy is measured rather than claimed: each month a random sample of the coder's claims goes to a certified coder, and every measure above comes from the log alone.
Membership and concierge billing
Plans carry a monthly fee, an annual discount and family pricing with each additional adult and child priced and capped. Enrollment caps apply per clinician, with a family counting each person against the panel. Employer contracts run at a per-employee-per-month price on one invoice to the employer, and the employee is never charged. Membership revenue is kept apart from insurance revenue by stream, and a revenue-split report shows them side by side, which matters to an owner deciding whether the membership line is carrying itself.
Direct primary care fees up to $150 a month for an individual and $300 for a family have been compatible with health savings accounts since January 1, 2026. The concierge skin prices against that line and produces the superbill a member files themselves, the one document a cash-pay practice cannot do without and the one most cash-pay systems leave to a spreadsheet.
Cash billing, with a complete receivables system
A cost estimate is called a treatment plan and is signed before the work begins, which keeps the conversation about care rather than about a bill arriving later. A plan is itemized from the price list for one patient or for several patients of one account, each line carrying a low and a high quantity so the plan shows a range. The signature is drawn by hand, and the signed text is fixed by a cryptographic fingerprint and written into the record. A deposit rule takes a percentage of the low estimate, a percentage of the high, a fixed amount or nothing, held as a liability until an invoice uses it. Charges after signing count against the chosen option's high estimate, and a charge that would exceed it is refused with the amounts shown, unless a clinician records that emergency care made a new signature impossible. In the human edition every treatment plan carries the No Surprises Act good faith estimate, with its federal timing rules and the patient's right to dispute a bill at least $400 above it.
The invoice is built at checkout for one account, each patient's lines grouped under that patient's name, showing coverage, discounts, tax, deposits applied, payments and the balance, and it can be split between accounts by percentage. Before checkout a missed-charge check compares what the record shows was done against what was charged.
Payments arrive as cash or check against an open drawer, by card through the processor with a token from a terminal or a hosted field so Arca never sees a card number, by bank transfer, or through a financing lender whose payment and fee are recorded separately. Each station opens one drawer with a counted float, and closing records the over or short by person. One end-of-day close reports production by class, collections by method, deposits, discounts, write-offs and tax collected. Refunds and write-offs each take a reason, and above the clinic's limit a manager or the practice owner approves, never the person who asked.
Every amount in the ledger, the receivables, the payables, the drawer and the price list is held and added as an exact decimal to the cent. A binary floating point number cannot hold most cent values exactly, so a balance accumulated in floats drifts by fractions of a cent and cannot be audited to the penny. One rounding rule applies: half up to the cent, once, on the line. A test fails if any stored ledger entry holds a float.
What the clinic sees, and who sees it
Role decides. Billers, the practice manager and the owner see the clinic's finances and its books. The front desk works checkout, every staff role sees an account's balance and its treatment plans, clinicians see what a visit costs the patient and nothing further, and a patient sees only their own statement.
A practice manager asks the analyst agent a question and gets an answer from that clinic's own data with the query shown and the numbers traceable. Why did collections drop in August breaks the change into its causes rather than returning a chart. Every read, write, denial and export of the ledger lands on a hash-chained access log anchored outside the database. There is no method in the ledger that takes another clinic's identifier and none that returns rows for research, and a clinic can export its whole ledger at any time.
The split between the two kinds of money record follows the line the law draws. Patient accounts name people and carry diagnosis and procedure codes, so they are protected health information and stay inside the clinic's ledger with the chart. The clinic's books need totals only, so each clinic's books are its own entity in the financial record, receiving one balanced journal entry per day. An entry is refused unless debits equal credits on accounts in the clinic's chart with no patient name anywhere in it, and a test posts a day and checks that no name reaches the books.
ARCA
The business side of the practice
A clinic is a business as well as a care setting, and it buys that half from somewhere else. Arca runs it: the people store, hiring and human resources, payroll, benefits, compliance, asset and supply tracking, bookkeeping with the books kept as a clinic entity in the financial record, and every tax return built line by line from those books for the clinic's own CPA to review and approve.
The products a clinic buys because its record system would not
A ten-clinician practice typically runs a payroll service, a human resources platform, a benefits administrator, a time clock, a compliance training vendor, a bookkeeping package, an asset register in a spreadsheet, and an outside bookkeeper to reconcile it all before the CPA sees it. Each holds the same people, the same dates and the same dollars, and each charges for the privilege. Each also takes a piece of the practice's operations: the practice cannot answer a question about its own staffing cost without exporting from three places and hoping the exports agree.
Arca runs that side of the clinic on the same record, under the same keys, with the same guide. The clinic stops buying separate products for payroll, human resources, benefits, compliance, bookkeeping and assets, and stops keeping the same facts in five places. The money that saves is the clearest commercial argument in the product, because it is money the practice is already spending and can stop spending the day it switches.
| What a practice buys today | What Arca runs instead |
|---|---|
| Payroll service | Gross to net, withholding by federal and state rules, a direct deposit file the clinic's own bank sends, tax deposits and filings |
| Human resources platform and time clock | Hiring, onboarding with work eligibility verification, the employee file, the time clock with federal and state overtime, paid leave, staff scheduling, reviews, separation |
| Benefits administrator | Plans, enrollment, the enrollment file to each carrier, pre-tax deductions, continuation coverage, federal reporting, premium reconciliation |
| Compliance training vendor and license tracker | The privacy and security program, incidents and breaches, training, safety logs, licenses with a hard stop on a lapsed one, credentialing, monthly exclusion screening, the compliance calendar |
| Asset register and supply ordering | The equipment register shared with the books, maintenance logs, supplies with lots and expiration, distributor ordering, the three-way match on vendor bills |
| Bookkeeping package and an outside bookkeeper | The clinic's own double-entry books as an entity in the financial record, fed one balanced journal entry per source event, with payables, fixed assets and the month-end close |
| Tax preparation from exported spreadsheets | Every return built line by line from the books, with each figure's basis named, for the clinic's CPA to review and approve |
The people store
One employee record per person, held in four sections each sealed under its own clinic key, because the sections are not equally sensitive and should not be equally readable. Each section carries its own retention date, with nightly destruction of what has passed it and a legal hold that suspends destruction for a named matter. The same record carries the role that decides what each person can do, the license types and their expiration dates, and the credential enrollment, and a lapsed license is a hard stop rather than a reminder.
Human resources, payroll and benefits
Hiring runs from the open position through onboarding, including work eligibility verification, into the employee file, and the time clock computes federal and state overtime, including California's daily rules, feeding paid leave accrual and staff scheduling. Payroll runs gross to net with withholding from the federal tables and each state's rules, produces a direct deposit file in the standard banking format that the clinic's own bank sends, and prepares the tax deposits and filings. Nothing is transmitted on the clinic's behalf, which keeps Arca out of the money transmission business. Two parallel pay periods run against the clinic's existing service before it switches, so the clinic sees its own numbers reproduced before it depends on them. Benefits cover plans and enrollment, the standard enrollment file to each carrier, pre-tax deductions, continuation coverage, federal health coverage reporting, and premium reconciliation. Each pay run, benefit premium and tax deposit becomes one balanced journal entry in the books, so the books are current without anyone copying anything.
Compliance, where the practice carries the liability
Compliance here is a register rather than a library of videos. It holds the privacy and security program with its named officers, incidents and breaches with their investigation and notification timing, training with completion by person, workplace safety logs, licenses, credentialing, monthly screening of staff against the federal exclusion lists, and one calendar of every recurring obligation with its owner and due date. Exclusion screening is the item worth naming: a practice that bills federal programs for services involving an excluded individual is exposed whether or not it knew, and most check once at hire. Here it runs every month, and a hit stops before it becomes a repayment.
Assets, supplies and purchasing
The equipment register is shared with the books rather than kept beside them, so an asset bought, depreciated or retired produces its journal entry from the same record that holds its maintenance history. Supplies carry lots and expiration dates, which is what makes a dispense traceable and an expired item refusable rather than merely flagged.
Purchasing runs against distributor catalogs with their own prices and item numbers, with the three-way match against the order and the receipt before a bill is approved. Payables cover capture from a photograph, approval by role, payment scheduling, contractor reporting at year end, returns and vendor credits, rebates, payment files and statement reconciliation. One item master serves the stock room, the clinic formulary and the price list, so a drug has one cost, one shelf quantity and one price rather than three records that drift.
The books, as the clinic's own entity
Each clinic's books are an entity of its own in the financial record, with a chart of accounts for medical practices, its own chain of events, profit and loss by month, location and provider, a balance sheet, direct cash flow, payables, fixed assets and the month-end close. Every dollar fact the business side produces becomes one balanced journal entry: a pay run, a tax deposit, a benefit premium, a vendor bill, an asset bought or retired, and the day's revenue summary.
The clinic's books hold no patient identity, by construction rather than by policy. An entry is refused unless it has at least two lines, each a positive amount on an account in the clinic's chart, with debits equal to credits and no patient or employee name in its memo or references. A line may carry the rendering clinician's staff identifier and the location code, and nothing more. The entry waits sealed in the clinic's outbox until the books accept it, and delivery is idempotent on its source, so a day posted again after late rows posts only the difference.
Audit walls separate the clinic from the owner's personal entity, a management company, a real estate entity and any other the owner holds, and money between them crosses only as a matched pair. A practice owner with a building, a management company and a clinic can see each audited on its own and still see the figures tie together, the arrangement most practice owners actually have and which no clinic software addresses.
Tax returns, prepared for the CPA
Every return the clinic owes is built line by line from the books, with the basis of each figure stated beside it, and handed to the clinic's own CPA to review, adjust and approve. Arca does not file and does not advise: the return is prepared work product, which is what a CPA bills a practice to assemble today from exported spreadsheets. Each figure that depends on the tax year is looked up for the year being prepared, and nothing is extrapolated. A filed return can be rebuilt exactly by replaying the record to the filing date, and an amendment is a new event pointing at the original.
The clinic stops paying six vendors to hold the same people, the same dates and the same dollars.
ARCA
Specialties, imaging centers and surgery centers
One code base serves every outpatient clinic type. A skin turns the core system into a specialty system by configuration, and imaging centers and surgery centers, which other vendors treat as separate products, run as two more skins over the same scheduling, charting, billing and patient portal.
What a skin is, and why forking is the thing to avoid
A skin is a configuration pack. It sets the note templates, flowsheets, rating scales, body diagrams, order sets, billing rules, device connections, schedule patterns and the widgets on each role's home screen. The core code is the same for every clinic type, and a multispecialty group runs several at once with each clinician seeing their own. Skins are versioned, so a coding or rule change reaches every clinic on that skin at once, and each clinic keeps its own adjustments in a layer above the pack. Each skin has a named coding and compliance owner and is updated on a set calendar, because code sets, payer rules and state laws change every year.
The alternative, which the field has chosen repeatedly, is to fork. A vendor sells a dermatology product and an orthopedics product, and the two drift until a fix in one has to be made twice. That is why one specialty vendor covers twelve procedural specialties with no primary care, no wellness and no imaging center billing, while the cash-pay systems have no specialty depth. Configuration instead of forking is what lets a multispecialty group run primary care, pediatrics, orthopedics, women's health, chiropractic, wellness, concierge, urgent care and imaging on one patient record.
A widget is a panel a clinic places on a home screen or inside the chart: the day's schedule, the waiting room, results to review, an infusion chair board, an operating room board, care-gap lists, today's collections. Each skin ships a default set, and every widget can be asked about by voice, so a nurse looking at a board can say why is chair 4 red rather than hunting for the underlying screen.
The specialty list
| Clinic type | What the skin sets | Billing rules it carries |
|---|---|---|
| General practice and primary care | Preventive and chronic templates, a care-gap engine, depression, anxiety and fall-risk screens; electrocardiograph, spirometry, waived analyzers, home blood pressure and glucose feeds | The complexity add-on, modifier rules with preventive visits, transitional care, risk adjustment coding |
| Internal medicine | Care plans, the annual wellness visit, a remote monitoring hub counting days of data per month | Care management time logs, advanced primary care management codes, the 2026 remote monitoring codes |
| Pediatrics | Growth charts with corrected age, well-child templates, developmental screens; two-way immunization registry, vaccine refrigerator loggers | Vaccine administration codes, federal vaccine program eligibility and separate stock, developmental screening codes |
| Orthopedics | Joint exams with laterality, fracture and operative notes, outcome instruments; image viewer, in-office radiograph worklist | Global surgical periods with post-operative tracking, split-care modifiers, bracing, injection waste reporting |
| ENT with allergy | Audiograms, tympanograms, allergy reaction grids, vial mixing and labels; audiometer, endoscopy capture | Audiology modifier rules, allergy testing and immunotherapy codes |
| OB/GYN | Prenatal flowsheet, due-date calculator with redating rules, prompts by gestational age, ultrasound report import | Global maternity billing with an episode ledger that survives an insurance change mid-pregnancy |
| Oncology | Regimen library, surface-area dosing, cumulative doses, toxicity grading, staging; infusion pumps, tumor registry export | The infusion code hierarchy, drug units with waste reporting, the federal oncology model's reporting |
| Wellness, integrative and IV therapy | Good-faith exam before treatment, infusion protocols, injectables by site, before-and-after photographs, retail point of sale | Packages, memberships, gift cards, retail inventory and sales tax |
| Concierge and direct primary care | Long-visit templates, the annual comprehensive exam, wearables through the patient's own application | Memberships, health-savings-compatible pricing, cash price lists, superbills |
| Urgent care | Triage acuity, fast templates, occupational medicine, waived testing, point-of-care radiography with overreads | After-hours codes by payer contract, employer billing, workers' compensation |
| Behavioral health integration | Depression and anxiety scores over time, suicide risk screening, a collaborative care registry | Collaborative care codes with the 2026 add-ons, substance use consent and labels |
Five skins run today on invented clinics: primary care, concierge and direct primary care, wellness and integrative, imaging center and surgery center, with three more in the veterinary edition. The rest are scoped, and the order they reach real clinics is set by which outside gates each group needs.
Imaging centers
An imaging exam runs from order to final report through one path. A radiologist sets the protocol, the contrast, the agent and the volume, and a with-contrast code cannot be protocoled without contrast or the reverse. Screening rules follow the professional contrast guidance, owned in a deployment by the clinic's medical director: contrast with a kidney function below the threshold holds, the band above it is a caution, and a recent result is required over a given age or with kidney disease or diabetes. A severe prior reaction holds, an unsafe implant holds, and pregnancy holds for ionizing radiation while unknown leaves the screening incomplete rather than passing it. A hold is overridden only by a clinician, with a reason, recorded.
Scheduling then books the scanner, the room, the equipment and the technologist together, issues the accession number and pushes the worklist entry to the imaging platform. The scanner reports progress back, and the final report writes its impression into the patient's record as an imaging result.
Billing follows the center's own arrangement rather than one assumption. A split arrangement bills an institutional claim at exam completion, with the revenue code for the modality and contrast on its own line, and a professional claim for the read at the final report. A technical-and-professional arrangement bills both as professional claims, and a global arrangement bills one claim at the report. That choice is a setting, because a center that cannot bill the way its contracts are written cannot use the system.
Surgery centers
A surgical case runs from scheduling through discharge with the checks enforced rather than documented after the fact. Scheduling books the room, the surgeon, the anesthesia provider and the equipment, and attaches the surgeon's preference card as the pick list. The preoperative checklist covers identity, consents, the history and physical within 30 days, the site marked when the procedure has a side, fasting, allergies, the anesthesia assessment and antibiotic timing, and the case cannot enter the room until every required item is done.
At the time-out the patient, the procedure code and the side must match the case and the site mark, and otherwise the time-out stops and the stop is recorded. No incision is recorded before the time-out. The intraoperative and anesthesia record captures in-room, anesthesia start, incision, closure, anesthesia end and out-of-room in order, with staff and roles, the anesthesia type, airway, physical status class and vital signs.
Implants are logged by scanning their device identifier, parsed from the standard barcode with its check digit verified. The device must be in the center's catalog, an expired device is refused, and a serial already logged is refused. Case costing runs on supplies and implants at unit cost, wasted items included, plus operating room minutes at the room's rate, the number an ambulatory surgery center lives or dies on and which most compute in a spreadsheet a month later.
Discharge requires the recovery scores at or above their thresholds, an escort present and instructions given, and is refused with the reasons otherwise. At discharge the facility claim, the surgeon's professional claim and the anesthesia claim are built together. The quality counts the federal program requires, for burns, falls, wrong-site events and hospital transfers, are kept with their denominators and rates.
Preference cards learn. After each finished case the last five cases for that surgeon and procedure are compared with the card: an open item used in none of them is removed, a quantity used in at least 60 percent of cases is changed when the median differs, and an item used in at least 60 percent is added. The suggestion waits for the card's surgeon, a nurse, a manager or an owner to approve, approval raises the card's version, and a suggestion made against an older version is refused.
A center's quality numbers come out of the record that enforced the checks, not out of a spreadsheet written afterward.
ARCA
Arca against the field
The ambulatory record market is crowded, well funded and already shipping AI. Arca was compared against more than 25 clinic systems and the leading AI notetakers, sorted into where it leads, where it must simply match the field, and where it was behind. Where a fact could not be verified in a public source, the comparison says so, and so does what follows.
Who holds the market
By installed ambulatory sites, Epic leads with 19.5 percent, followed by eClinicalWorks at 11.9, athenahealth at 6.9, Oracle Cerner at 5.4, NextGen at 4.2, ModMed at 3.6, and Veradigm and Practice Fusion at 3.1 each (Definitive Healthcare, accessed October 2025). In the 2026 Best in KLAS awards Epic and athenahealth split the ambulatory segments, and Elation won the small-practice segment for a second year.
Epic holds most large US health systems and about one in five outpatient sites. In December 2025 the State of Texas sued Epic, alleging that it blocks competitors and overcharges. That suit is the clearest public signal that the dominant position is being contested on conduct rather than on features, and it marks the segment where Arca does not compete first.
What it costs, and what it does badly
| Competitor | Owner and position | Price, published or reported | AI | Satisfaction |
|---|---|---|---|---|
| Epic | Private, employee-controlled; health systems and large groups | Not published; its small-practice product reported at about $2,000 a month plus $10,000 setup | Native AI charting from February 2026, a patient agent, a revenue cycle agent | Best in KLAS 2026; 2025 B-minus, 84% |
| Oracle Health | Oracle, after the $28B Cerner purchase; health systems | Not published | A voice-first certified record, a clinical agent with coding | No ambulatory score found |
| athenahealth | Bain Capital and Hellman & Friedman, $17B in 2022; independent and mid-size groups | A percentage of collections; about $140 per provider cited by third parties, not verified | Ambient notes at no cost from August 2026, an inbox agent in testing | Best in KLAS 2026; 2025 B-minus, 77% |
| eClinicalWorks | Private, founder-owned; small to large practices | Median buyer report $18,750 a year; its scribe $149 to $199 per user a month | A scribe, a contact center agent | 2025 C-minus, 62% |
| Veradigm and Practice Fusion | Over the counter after the 2024 delisting; a 2025 sale review ended without a deal | Practice Fusion listed at $199 a month | A billing agent | 2025 D, 43% |
| ModMed and Elation | Clearlake majority at $5.3B; venture-backed primary care | Not published; quote only | A scribe and a scheduling agent; note assistance and agentic billing | Elation Best in KLAS 2026, 1 to 10 physicians |
| Hint, Akute, Atlas.md | Private; direct primary care, concierge, cash-pay | $290 to $770 a month; $300 to $550; about $100 to $300 per provider | Thin, or not verified | Not verified |
| Healthie, Practice Better, CharmHealth, Canvas | Private and venture-backed; wellness, integrative, virtual-first | $19 to $149 a month; $0 to $155; $200 per provider; $4,000 a month for 1,000 patients | Charting add-ons; one claim agent, certified | Not verified |
Two things read off that table. The low end of the market, where Arca enters, is priced between about $100 and $770 a month and its AI is a bolted-on scribe. And the published satisfaction scores run from a B-minus down to a D. Nobody is winning this field on the experience of the software.
The AI scribe has been commoditized
In the past year the AI-drafted note stopped being a product. Epic shipped native charting in February 2026, athenahealth made its ambient notes a no-cost part of its platform in August 2026, Oracle launched a voice-first certified record in late 2025, and Doximity gives a scribe to verified US clinicians for free. Among the standalone vendors, Abridge was valued at $5.3 billion in June 2025, Microsoft's product is priced at $3.00 per ambient note, and Heidi raised $340 million in September 2026 at 2.8 million visits a week.
The price of a standalone scribe has collapsed toward zero, and Arca cannot charge for the note. The contest has moved to agents that code, bill, schedule and answer patients, and every major vendor has shipped one. What none of them has shipped, as far as any public material shows, is one spoken guide across clinical and front-office work with a person approving each action, over a record that weighs the source of every value in it.
| Where Arca stands | What that means |
|---|---|
| Ahead: voice runs the practice, not just the note | Competitors apply voice to documentation and chart lookup, each agent automating one slice. No surveyed vendor claims one guide across the whole day with approval at every step |
| Ahead: a learned trust weight on every value | No competitor found stores one. The nearest idea links note text to a recording |
| Ahead: one code base across specialty and cash-pay clinics | The specialty vendor has no primary care, wellness or imaging center billing; cash-pay systems lack specialty depth; the one system covering everything is priced out of small practices |
| Ahead: each clinic's finances under its own key | A 2026 breach at one European health IT company, put at 11 to 15 million patient records, makes security design a live buying question |
| Must match: AI notes, voice navigation, AI coding, AI phone lines, a patient application, tablet clients | All present in the field, and the note is given away by two incumbents. Arca shows breadth rather than novelty, reporting voice coverage per release |
Outside the United States
National primary care markets are held by long-standing incumbents, and the money is moving: EMIS at about 55 percent of English general practices and TPP SystmOne at about 35, with TPG buying Optum UK for $400 million in March 2026; Doctolib and Cegedim in France; CompuGroup Medical in Germany, taken private in 2025; Best Practice and MedicalDirector in Australia with no native AI found; TELUS Health and Accuro in Canada. Entry abroad takes private and cash-pay clinics first and public-sector systems last.
The certification path, and what it gates
No federal law requires a record system to be certified. Clinics that bill Medicare need certified software to report in the federal quality program, and most will not buy without it. That fact sets the sequence: Arca launches with concierge, direct primary care, wellness and cash-pay practices, then certifies after that pilot and after insurance billing.
The program's baseline requires the federal core data set from January 1, 2026, transparency for decision support so a user can see the attributes behind each intervention, a standard patient and population interface with third-party sign-in, full export of one patient's data and of a whole clinic's, electronic prescribing, and prior authorization interfaces from October 2025. A proposed rule from December 2025 would remove 34 of the 60 criteria; its final status was not verified, and the list is set once it is final.
Certification also carries a cost that is not a fee. A certified developer becomes subject to the federal information blocking rules, with civil penalties of up to $1 million per violation, and federal investigations were announced in February 2026. Arca is built to that standard already, which is why results release on finalization by default. Network participation runs through one qualified health information network, an exchange that passed one billion records by June 2026 across eleven networks.
What is still to be certified
Arca has not entered certification testing and no certification has been granted. What exists is the software certification would test, built and tested on invented data. The outside work that remains:
- Certification test procedures at an authorized test laboratory, across orders, care coordination, export, decision support, quality measures, privacy and security, and patient access.
- Validation of the clinical document output and the prior authorization bundle against the federal validator and its profiles.
- The third-party interface test kit, and the consent screen for outside applications, which does not exist yet.
- Prescribing network certification, and a third-party audit of the controlled substance signing path.
- Registration with each state immunization registry, and a participation agreement with a qualified health information network.
- Formal user-centered design testing with real users, a quality management system, accessibility conformance.
- An outside penetration test, a privacy and security risk analysis with named officers, and the security attestations clinics ask for.
- A named practicing specialist and coding owner for each specialty content pack.
- No outside satisfaction rating of any kind, since ratings follow live clinics.
ARCA
The Arca market
Arca sells into two published market segments worth $47.5 billion to $56.3 billion a year worldwide, and into a count of institutions that is harder to argue with than any market estimate. The entry market is the one that needs no federal certification to buy, and the arithmetic that matters is the number of practices multiplied by a price a small practice already pays.
The market Arca sells into
Health records and practice management together, worldwide.
The burden Arca is built against
A primary care physician spends about 36 minutes in the record for every 30-minute visit (AMA). In academic primary care from 2019 to 2023, time in the inbox grew 24 percent and time spent on orders grew 59 percent, and after-hours work has not fallen even where scribes are in use (Annals of Family Medicine, AMA). The incumbents added scribes to systems built around forms and billing screens, and the work moved instead of shrinking.
Ambient scribes save one to two minutes a visit in controlled studies. The note is the small part of the burden, and the parts that grew fastest are the parts no scribe touches. That is the whole positioning: the orders, the coding, the inbox, the front desk and the prior authorization work are where the agent team is aimed.
Sizing, with the sources named
| Segment | Global, per year | United States |
|---|---|---|
| Electronic health records | $30.3B to $37.5B | $9.4B to $15.0B |
| Practice management | $17.2B to $18.8B | $5.1B |
| Arca's two segments together | $47.5B to $56.3B | $14.5B to $20.1B |
| Hospital information systems, adjacent | $63.8B | Not separated at this scope |
Each line is a range between the low and high published estimate because research firms define these markets differently. The figures are pooled from published estimates by Fortune Business Insights, Grand View Research, Precedence Research, Mordor Intelligence, Towards Healthcare, MarketsandMarkets, Global Market Insights, Expert Market Research, Persistence Market Research, Nova One Advisor, IMARC and Market Research Future, rather than from one commissioned study. Anyone needing a figure attributed should take it from the publishing firm's own report. The two segments overlap in practice, since a clinic buys one system that does both, and the hospital line is adjacent rather than addressable until outpatient certification is done.
The ambient scribe market, the segment the field is currently fighting over, was about $600 million in 2025, led by Microsoft at 33 percent and Abridge at 30 percent. It is roughly one percent of the two segments above, which is the size of the prize the incumbents are giving away to protect the larger one.
Counting customers instead of dollars
Market sizes are estimates built on other estimates. Counts of institutions are harder to argue with.
| Who | How many | Source |
|---|---|---|
| US physician group practices | 395,000 | Definitive Healthcare |
| US imaging centers | 15,000 | Definitive Healthcare |
| US urgent care centers | 11,000 | Definitive Healthcare |
| US ambulatory surgery centers | 6,436 | MedPAC, March 2026 |
| US direct primary care practices | 2,700 and more | DPC Frontier via Atlas.md |
| Hospitals in the Gulf states | 882 | GCC Statistical Centre |
The arithmetic
One assumption shows the scale. If Arca averaged $6,000 a year per practice, which is about $500 a month and inside what small practices pay today, the 395,000 US group practices alone would be a $2.4 billion yearly market for Arca.
The assumption is checkable against what the field charges. Direct primary care and cash-pay systems run about $100 to $770 a month per clinic or clinician. Wellness and virtual-care systems start under $50 a month. The eClinicalWorks median buyer report is $18,750 a year. A standalone scribe adds $149 to $199 per user a month on top of whichever of those a practice already pays. Against that, $500 a month for the record, practice management, the claims cycle, the patient application, the front office and the business side of the clinic is not an aggressive figure, and the imaging center assumption of $24,000 a year is a fraction of what a center spends on its archive and reporting stack today.
The $6,000 figure is an assumption in the model and not a quotation or a signed price. Prices are set with the first clinics during the pilot. Changing the assumption changes the arithmetic, and it is written this way so a reader can change it.
The entry market
The first clinics are cash-pay primary care, direct primary care, concierge, wellness and integrative practices. Three things make that the entry point rather than a compromise: they carry the lowest certification burden, since they do not bill Medicare and therefore need no certified software to buy; their current vendors ship the thinnest AI in the field, mostly a bolted-on scribe; and their buyers value the patient application and the lifetime record more than a hospital does, because the relationship with the patient is the product they sell.
The second group is small practices on the lowest-rated incumbents, which report the worst satisfaction in the field, 43 to 62 percent, on vendors in unstable ownership. The third is multispecialty groups and management services organizations, which is where the single code base is worth the most, because they run several specialties on several vendors today and no incumbent covers primary care, pediatrics, orthopedics, women's health, chiropractic, wellness, concierge, urgent care and imaging on one patient record. Health-system-owned clinics come last, after certification, network exchange and outside ratings.
Distribution into that third group is not a cold-call business. The founder has run, managed or consulted for more than 100 medical, dental and veterinary clinics and founded two management services organizations, which is the buyer this product is aimed at, approached by someone who has sat in that chair.
International markets
The order abroad is set by how much regulatory work stands between the software and a paying clinic, and each market's requirements are known before entry rather than discovered during it.
| Order | Market | What entry requires |
|---|---|---|
| First | United Kingdom private clinics, Australia, Canada | In the UK: the digital technology assessment, clinical safety standards with a named safety officer, the data security toolkit, cyber certification, UK data protection law, and the medicines regulator for any device function. In Australia: the privacy act, national record conformance, the benefits schedule for billing, healthcare identifiers and electronic prescribing conformance. In Canada: provincial privacy law such as Ontario's, provincial record certification and billing, and Canadian hosting |
| Second | Private hospital groups in the Gulf states, India and Latin America | In the Gulf: the national exchanges with in-country storage, and Saudi Arabia's platform and data localization. In India: the national digital health identity and exchange, and the 2023 data protection act. In Latin America: Brazil's data protection law and its billing standard, and Mexico's record standard |
| Third | The European Union | The new European health data space conformity process, once it is in place from March 2027 |
| Last | National public-sector primary care | Incumbents hold decades-long positions, procurement runs on a different clock, and the competition abroad is national in scope |
The reason private and cash-pay clinics come first abroad is the same reason they come first at home. National primary care markets are held: EMIS at about 55 percent of English general practices and TPP SystmOne at about 35, Doctolib and Cegedim in France, CompuGroup Medical in Germany, Best Practice and MedicalDirector in Australia, TELUS Health and Accuro in Canada. Private hospital groups in the Gulf states, India and Latin America are open on price, because a full incumbent installation is out of reach there, and the Gulf alone counts 882 hospitals.
Market sizes are estimates built on other estimates. Three hundred and ninety-five thousand practices is a count.

VETERINARY
The veterinary market
Veterinary practice software is a separate business from human clinic software, with its own buyers, its own incumbents and almost none of the regulatory weight. It is also the one market where a complete clinic system can be shipped into revenue without waiting on federal certification, because there is no insurance claim to certify against.
The veterinary software market
Veterinary practice management software, worldwide.
The size of the thing being sold into
The American Pet Products Association puts total US spending on pets at $158 billion in 2025, with veterinary care and product sales at $41.0 billion of it, and projects veterinary care and product sales at $42.4 billion in 2026. IBISWorld estimates US veterinary services industry revenue at $72.6 billion in 2026, a wider definition taking in laboratories and other veterinary businesses alongside clinics.
The practice count depends on what is counted. The American Veterinary Medical Association's 2025 report on the economic state of the profession carries the US Census Bureau figure of about 34,000 veterinary establishments in 2022, up 18.5 percent since 2009. IBISWorld counts 57,920 veterinary services businesses in 2026, a figure that includes sole practitioners, mobile practices and non-clinical veterinary businesses. Either way the number of buying locations runs to the tens of thousands, every one of them running a record, a schedule, a pharmacy and a cash register.
Demand under those practices is not a projection. The AVMA's 2024 household survey counted 89.7 million dogs and 73.8 million cats, with 59.8 million households keeping a dog (45.5 percent of US households) and 42.1 million keeping a cat (32.1 percent). It put average yearly veterinary spending at $580 for a dog household and $433 for a cat household, and the average cost of a visit at $147 in 2024.
What practices pay, and who sells it to them
Two companies hold most of the installed base, and both sell the practice something else as their main business. IDEXX Laboratories, a diagnostics company, owns Cornerstone and bought the cloud system ezyVet in 2021. Covetrus, a distribution and pharmacy company, carries AVImark, Impromed, Pulse and eVetPractice; Clayton, Dubilier & Rice and TPG took Covetrus private in 2022 at an enterprise value of about $4 billion, as the buyers announced at the time. The practice system in both cases is the hook that keeps a clinic buying laboratory work or drugs, which sets how fast it gets rebuilt.
| Incumbent | Held by | What it is |
|---|---|---|
| Cornerstone | IDEXX Laboratories | The long-standing installed system in general practice |
| ezyVet | IDEXX Laboratories, acquired 2021 | The cloud system IDEXX bought to answer the cloud entrants |
| AVImark, Impromed | Covetrus | Older practice systems carried alongside distribution and pharmacy |
| Pulse, eVetPractice | Covetrus | The cloud line of the same group |
| Shepherd, Digitail, Vetspire, Instinct | Independent, venture-backed | Newer systems, each strong in one part of the practice |
Published prices put the category in the hundreds of dollars a month. Costbench's 2026 comparison of ten veterinary systems lists monthly prices from $49 to $600, averaging $227, with ezyVet at $260.50, Covetrus Pulse from $200 to $600, AVImark from $150 to $400 and Digitail from $149 to $300. The IMARC Group puts the global veterinary software market at $616.0 million in 2025, reaching $984.7 million by 2034. Tens of thousands of practices paying a few hundred dollars a month is the arithmetic that makes the incumbents comfortable, and it also means a switching decision is made on what the software does rather than on a contract large enough to need a committee.
Why the regulatory path is short
The reason human clinic software takes years to ship is the insurance claim. A system that bills Medicare must be certified, which means a published criteria list, a testing laboratory, a surveillance program and a release cadence set by a federal agency. Veterinary medicine has no equivalent: no claim to submit, no payer to certify against, no certification program to enter.
That follows from how the care is paid for. The North American Pet Health Insurance Association reports about 7 million insured pets in the United States in 2025 against $5.7 billion of in-force gross written premium, which is 4.3 percent of US pets. More than 95 percent of American veterinary medicine is paid for at the counter by the person who brought the animal.
What remains is real and is built: FDA rules on extra-label drug use, the DEA's duties for a veterinarian holding controlled drugs, state prescription monitoring where a state requires it of veterinarians, the rabies certificate and its reporting to the local licensing authority, and state sales and use tax on drugs and retail goods. Each is carried as a rule per state and checked before a clinic goes live there. None of them is a gate that stops a sale.
The crossing only runs one way
A clinic system built for human medicine already contains everything a veterinary practice needs: scheduling, the chart, orders, results, imaging, surgery and anesthesia, inventory, a pharmacy, payroll, the books, a patient application. Making it veterinary means changing the subject of the record from a person to an animal with a species, adding veterinary references and vocabulary, and switching the insurance modules off.
Going the other way means building the parts a veterinary system never needed: eligibility, claims, remittance, denials, prior authorization, coding audits, payer rules, the clearinghouse connection, and the certification that lets any of it bill.
Adding veterinary medicine to a human clinic system is a configuration. Adding human medicine to a veterinary system is a second company.
Why this company and this market
Revenue before certification
Veterinary practices need no certified software to buy, which puts a complete clinic system in front of paying customers while the human editions work through certification, network exchange and the payer connections.
The founder has run these clinics
Greg has founded and run veterinary clinics, and has run, managed or consulted for more than 100 medical, dental and veterinary practices. The decisions in the veterinary edition came from that chair rather than from a requirements interview.
The billing problem is the simple one
A practice billing the person in front of it needs receivables, deposits, a drawer and collections done exactly. It does not need a claims cycle, which is where most of the cost and all of the certification sits.
Buyers who standardize across locations
Brakke Consulting's estimate, reported by the American Animal Hospital Association, puts about 25 percent of general practices and about 75 percent of specialty practices under corporate ownership, holding close to half the market's revenue.
Consolidation is the distribution story. The AAHA report names NVA with more than 1,400 hospitals, Thrive Pet Healthcare with 380 clinics, and a planned merger of Mission Veterinary Partners and Southern Veterinary Partners covering 730 practices, against PitchBook's count of $51.6 billion of private equity investment in the sector. A group running hundreds of locations on a mix of inherited systems has the most to gain from one record, one pharmacy, one price list and one set of books, and is the only buyer able to move a thousand practices on a single decision. The remaining three quarters of general practices are independent, paying a few hundred dollars a month for a system whose vendor's real business is selling them laboratory work or drugs.

ARCA VET
Arca Vet
Arca Vet is the veterinary edition of the clinic system: one codebase with the clinic's edition set to veterinary. The Patient is the animal, the Client is the person or organization who brings it and pays, and the clinic bills Clients directly with no insurance claims.

One codebase, two editions
Every clinic is set up as one of two editions, human or veterinary. The setting sits on the clinic beside its skin, and a practice treating both people and animals runs two clinics. It can be changed only before the clinic's first patient is registered.
| The edition decides | Human edition | Arca Vet |
|---|---|---|
| The patient model | A person in the identity vault, with a responsible party for a minor or a dependent adult | An animal Patient with its signalment, paid for by one primary Client and any co-clients |
| Which modules answer | Every module | Every module except insurance. Eligibility, claims, remittance, denials and prior authorization answer that they are not available |
| References | Human drug reference, dose limits and laboratory ranges | Species and breed codes, per-species dose and laboratory ranges, veterinary views and forms |
| Words on the screen | Patient, responsible party, co-responsible party | Patient (the animal, by name), Client, co-client |
| Outside messages | The human prescriber element, the human species code in state reports | The veterinarian prescriber element, the veterinary species code, the imaging standard's animal attributes |
| The patient's own app | My Talisman | Pet Talisman |
The insurance refusal is enforced by a check on every request rather than by hiding buttons, so a route that should not exist in a veterinary clinic cannot be reached by guessing its address. Membership becomes the wellness plan, and every other module is shared outright.
The Client is the account, the Patients hang off it
The Client is a billing account: a person, or an organization such as a shelter, rescue or breeder. A Patient has one primary Client and may have co-clients, each with their own share of the bills, their own right to see the record and consent, and their own statement, and one account can pay for any number of Patients. Responsibility for bills is kept separate from access to the record, so a co-client's share of an invoice and a co-client's right to read the chart are two different grants. A Patient can move to a new Client after an adoption, with the earlier Client keeping their past invoices while losing access. The same account model carries the human edition's responsible party, the person financially responsible for a minor or a dependent adult, which came out of the veterinary work and went into the human edition.
The animal Patient
Each Patient is a record subject of kind animal, matched only against animals and never merged with a person. The evidence-weighted record works for an animal exactly as for a person: an in-house analyzer and a reference laboratory reporting the same value are weighed the same way.
What the animal is
Name, species and breed, sex and reproductive status, birth date with an estimated flag, color and markings. A male cannot be recorded spayed. A mixed breed is a predominant breed plus up to two others, or mixed with a size class.
Species and breed as codes, not free text
Eleven species load from the Vertebrate Breed Ontology, ten of them with their breeds, and a breed code from the wrong species is refused. Food species are flagged, because small-animal practices see pet pigs, goats and backyard poultry.
Microchip, rabies tag, license, tattoo
A microchip number is checked against its format, and a second Patient with the same chip is refused. The chip carries a blind index, so a scanner finds the Patient without opening every record in the clinic.
What gates the dose
A dated weight history, where a weight older than 30 days, or 7 days under one year, refuses a per-kilogram dose. Body condition, pain score, allergies, and genetic markers such as MDR1 as normal, carrier, affected or untested.
A handling caution is written as an instruction ("muzzle for nail trims"). Labels describing the animal's character are not used, because they travel between staff as judgments and change how an animal is treated. A death stops that Patient's reminders within the hour.
The vocabulary is not a label change
The word "owner" never describes a person's relationship to an animal: not on a screen, a form, a label, a certificate or a message, and not in the guide's speech. A test reads the veterinary patient screens and fails if the word appears. The staff role for the person who owns the practice shows as "Practice owner" in both editions, a role in a business and not a claim about an animal. In Pet Talisman the same person is the Parent and the animal is called by its name, and a cost estimate is a Treatment Plan in all three editions. Two outside formats print their own label and are left alone: the rabies certificate and a few state prescription forms print the issuing body's wording for the person's field, with the Client filled into it.
Veterinary clinical content
Directional terms are veterinary: cranial, caudal, rostral, dorsal, ventral, palmar, plantar. Radiographic views are named for the direction the beam passes and coded to the imaging standard's view list for animals, so the view a human chart calls AP is ventrodorsal or dorsoventral. Laboratory results carry a reference range for the species, and one with no range shows "no reference range for this species" rather than a human range, because the human range is a wrong answer rather than a missing one. Vaccines run from a per-species catalog the clinic adopts under its medical director, each vaccination recording the manufacturer, product, lot and expiration, site, route, who gave it and the next due date. A rabies vaccination prints the national certificate's every field with the veterinarian's signature.
A hospitalized Patient has a treatment sheet of scheduled treatments, medications, fluids, monitoring and feeding. Marking a treatment done records it and adds its charge, so nothing done in the hospital goes unbilled, the largest source of lost revenue in a practice with inpatients. Dental charting uses modified Triadan numbering per species, and the end-of-life record carries a euthanasia consent with aftercare choices, after which reminders stop and the account stays open for its balance.
The photo, and confirming the right patient
The photo appears on every screen that names a patient: the chart header, the schedule, check-in, the hospital board, the dispensing screen and checkout. Check-in asks for a new one when it is older than 12 months for an adult, 3 months for an animal under one year, or when the weight has moved 10 percent or more. Before a procedure, an anesthesia induction, a euthanasia, a transfusion, a dispense or a controlled-drug administration, the screen shows the photo with the name and asks the staff member to confirm. The veterinary edition also offers a microchip scan, and a chip belonging to another Patient or to none on file stops the step and is logged. The same photo is built into the human editions, taken only with the patient's consent, never used for automated face matching and never leaving the clinic. A study at a large academic medical center found fewer wrong-patient order errors when the record displayed the patient's photograph (Salmasian et al., JAMA Network Open, 2020).
What the edition shares rather than forks
A separate veterinary product would mean building the chart twice, the schedule twice, inventory twice and the books twice, then maintaining two of everything while they drift apart. Arca Vet shares the record, the identity vault, scheduling, orders and results, the imaging path, surgery and anesthesia, the item master, the formulary and pharmacy, cash billing, payables, payroll and the books.
The sharing runs both ways. The clinic formulary, the complete receivables and payables, the Treatment Plan, the patient photo and the perpetual controlled-substance log were built for the veterinary edition and landed in the human edition at once, because a wellness clinic dispensing from its own shelf needs every one of them.
The veterinary edition is the same product with the insurance taken out and the animal put in.
A clinic day, end to end
A full veterinary clinic day runs across the products together: Arca Vet, Pet Talisman, the Holarc record and Finarc's books. A Client registers two dogs with a co-client, and both are photographed at check-in. A dental Treatment Plan is signed with a deposit, and the controlled drug given under anesthesia is logged with a witnessed waste. The dose check refuses a drug that is dangerous to a Collie carrying the MDR1 mutation. Another drug comes from the first-expiring lot with its label, and the rabies certificate prints. The invoice splits between the two co-clients, one paying by card and one in cash, and the drawer closes even. The day posts to the clinic's books with no Client or Patient name in them. The co-client pays the rest from Pet Talisman, and both balances read zero in both applications. Every step passes on invented data.

ARCA VET
The clinic pharmacy and cash billing
Most veterinary practices run at least part of their own pharmacy, which makes dispensing a core function rather than a module. The formulary, the shelf and the price list are one list, and the money behind them is a complete receivables and payables system, because a practice that bills the person in front of it has nowhere else to send the bill.
Why the pharmacy is the center of a veterinary practice
A human clinic writes a prescription and the patient takes it somewhere else. A veterinary practice fills it at the counter, from stock it bought, under its own license, and charges for it on the same invoice as the exam. Drugs and preventives are a revenue class of the practice rather than a referral. A system that treats dispensing as an afterthought leaves the practice running its pharmacy in a spreadsheet, which is how controlled drugs go missing.
The formulary entry, the inventory item, the lot on the shelf and the price list entry are four views of one thing. A dispense runs across all four: the dose is checked, the quantity leaves the first-expiring lot, the label prints, the line joins the invoice, and the cost posts at lot cost.
The formulary and the dose check
The clinic formulary is the list of drugs the practice stocks or prescribes, curated by its medical director. Each entry holds the drug by established name and ingredients, the stocked item that supplies it, the price list entry that charges for it, and the label text. In the veterinary edition it also holds, per species, whether the drug is approved for that species, and a dose range in milligrams per kilogram with frequency, route and maximum dose.
Cautions attach by species, breed or genetic marker, each with its own action: warn, require a reason, or refuse unless the veterinarian overrides. Acetaminophen and permethrin are dangerous to cats; dogs and cats carrying the MDR1 mutation need reduced doses of several drugs, on the list Washington State University publishes. Each entry carries its controlled-substance schedule federally and in each state where the clinic operates, because some states schedule drugs the federal list does not.
| The dose check, on every order and dispense | What happens |
|---|---|
| Weight age | A weight older than the clinic's limit refuses the dose: 30 days, or 7 days under one year of age |
| No range for the species | Refused until the veterinarian enters a range and accepts the use as extra-label |
| Dose outside the range or above the maximum | Needs the veterinarian's reason, recorded with the order, and marks the use extra-label |
| Drug not approved for the species | Extra-label even inside the range |
| Genetic caution | Follows the Patient's tested status, with the untested action applying to an untested Patient of a named breed |
Dispensing, labels and stock
The veterinarian orders and a technician, nurse or the veterinarian fills from the formulary. The system refuses an inactive entry, an entry with no stocked item, and directions missing frequency, route or duration. It runs the dose check, requires the condition treated for any extra-label use, requires a photo confirmation of the Patient from the last sixty minutes and a second for a controlled drug, and takes the quantity from the first-expiring lots. The label carries every element the FDA requires for an extra-label dispense and the AVMA recommends for every label, down to the withdrawal time for a food species.
One item master sits behind all of it, each item linked to its formulary entry if it is a drug and its price list entry if it is sold. Every dispense, administration and retail sale relieves stock at lot cost when it happens, so a physical count then shows only true shrink. Part of a vial is charged as given, so half a milliliter is half the per-milliliter price. A vendor bill is matched to its purchase order and to what was received, and a quantity billed above what arrived stops at approval.
Controlled drugs
A perpetual log runs per container, recording every receipt, administration, dispense, transfer, waste, count, loss and disposal with the Patient, the amount, the person and the balance after. The balance can never go below zero, and waste is witnessed by a second person, never the one who wasted it.
A receipt opens a container with its lot, expiry and that location's DEA registration, which every controlled record then carries. The running balance is checked at every shift change, and a shortfall prepares the federal loss report straight from the log. An expired drug leaves the log only through a reverse distributor or a destruction recorded with a witness. States differ on whether a dispensing veterinarian must report to the state monitoring program, so the rule is carried per state. The same log serves the human editions.
A Client may ask for a written prescription instead of dispensing, which professional principles and most states require. Electronic prescriptions go out on the national prescription standard with the veterinarian prescriber element and the animal patient layout.
Cash billing, because the Client pays
Arca Vet bills Clients directly and sends no insurance claims, which makes receivables the whole revenue cycle rather than the remainder after a payer has been dealt with. A cost estimate is called a Treatment Plan, signed by the Client before the work begins. The name keeps the conversation about care.
A plan is itemized for one Patient or for several Patients of one Client, each line with a low and a high quantity so it shows a range, and it can offer options side by side such as medical management and surgery. The signature is drawn by hand with a typed name beside it, fixed by its fingerprint, and written into each Patient's record as a signed consent. A deposit rule runs per plan and is held as a liability until an invoice uses it. Charges after signing count against the chosen option's high estimate, the front desk is alerted at 80 percent, and a charge that would exceed it is refused, unless a clinician records that emergency care made a new signature impossible.
Charges wait per Patient from the front desk, dispensing, the hospital treatment sheet and the wellness settlement, and a missed-charge check before checkout compares what was done against what was charged. The invoice groups each Patient's lines under that Patient's name and photo.
Payment comes as cash or check against an open drawer, card through the processor with no card number reaching the system, bank transfer, patient financing with the lender's fee recorded separately, or a payment a Client's own pet insurer sends directly. Closing a drawer reports over or short by person. Open invoices age in 30-day buckets, and an account owing nothing is never sent to collections. A write-off or a refund needs a reason and approval by a manager or the practice owner, never by the person who asked. Discounts post against revenue so the books show what was given away.
Every amount in the ledger, the receivables, the payables, the drawer and the price list is held as an exact decimal to the cent, rounded half up once on the line. A binary floating point number cannot hold most cent values exactly, so a balance accumulated that way drifts by fractions of a cent and cannot be audited to the penny.
The human edition gets the same system
None of this is veterinary-only. A human clinic keeps its claims cycle, where only the patient's own share reaches the cash invoice, and gains the identical receivables and payables. A Treatment Plan for a self-pay patient also satisfies the federal good faith estimate, including the notice of the right to dispute a bill at least $400 above it. Direct primary care, concierge and wellness practices run almost entirely on cash billing and a dispensary, and the veterinary work makes that entry market properly served.
A veterinary practice has no payer to blame for a receivable. The money is either collected at the counter or it is not collected.

PET TALISMAN
Pet Talisman
Pet Talisman is the veterinary edition of the patient app: the same application with its edition set to veterinary. A Parent keeps each pet's lifetime record, reminders and clinic bills in one place, signs Treatment Plans and pays the clinic, shares proof of vaccination, and asks a guide that speaks about the pet by name, sends an emergency to care, and never gives a medicine dose for an animal.

The Parent and the pet
On screen the person is the Parent and the animal is called by its name. "Bella's vaccines," not "the patient's immunization history." The word "owner" is never used for a person's relationship to an animal, on a screen, in a message or in the guide's speech, and a test fails if it appears anywhere.
Each pet is a record subject of its own kind, created as an animal in the record core with its microchip in the crosswalk, matched only against animals and never merged with a person. Its name, species, breed, sex, reproductive status, birth date, microchip and photo are sealed field by field. A microchip already in the app cannot be added twice; the second household asks the pet's Parent for an invite. The human proxy rules do not apply to an animal: no under-18 end date, no adult grant, no legal representative, because none of them means anything for a dog.
When a pet dies the record stays, reminders stop the same hour, sharing links end, and the pet moves to a memorial section the Parent can open whenever they want. Nothing prompts the Parent about that pet again, because the alternative is a vaccine reminder three weeks after a euthanasia.
The lifetime record
The clinic's front desk gives the Client an eight-digit code. The Parent enters it in Pet Talisman, which links their sign-in to their Client account. Every call to the clinic then carries the Parent's own token, so the app reaches that account's pets, bills and records and nothing else.
| What the Parent has | What it holds |
|---|---|
| The pet's record | Weights, vaccines with the next due date, preventives, problems, allergies, and each signed visit with its home instructions |
| Laboratory results | Released results with the range for the pet's species, or "no reference range for this species." A human range is never shown for an animal |
| Medications | Each medicine the clinic dispensed, with the label's directions and course dates, and dose reminders in the Parent's time zone |
| The rabies certificate | Every field of the national certificate, to download or share |
| Invoices and Treatment Plans | Each invoice grouped by pet, with the Parent's own share where it is split between co-clients. Plans are reviewed, chosen, signed with a drawn signature, and the deposit paid there |
| Plans and payment | The wellness plan with what it covers this plan year, and the balance over 2 to 12 monthly card payments, the first charged at once |
The Parent books and cancels visits, fills in check-in forms with the pet's name in each question, messages the clinic in a thread per pet, asks for refills, and sends photos from home. A photo from home is marked as from the Parent; the identity photo used to confirm a patient before a procedure is always one the clinic took. Reminders run on vaccine and preventive due dates, the wellness plan and any recheck the veterinarian orders, and they stop for a deceased or transferred pet. Preventive care is where a practice loses money quietly and where an animal loses years, and a reminder in an app the Parent already opens to pay a bill is read.
Sharing, the lost-pet flyer and the insurance packet
A link for a groomer or a kennel
Proof of vaccination or the full record, for 1 to 365 days as the Parent sets, opened without signing in. The record is copied into the link and sealed, with the Parent's address left off. Links can be turned off and count their opens.
A flyer in one button
The pet's photo, name, description and microchip number with the contact the Parent chooses, ready to print. A Parent with a missing animal has no patience for a form.
A packet for the Parent's own insurer
The pet's details, the clinic's invoices and the vaccine record, as data and as a readable page, for the Parent's own claim. The clinic bills no insurer.
The guide, and what it will not do for an animal
The guide speaks about the pet by name and answers from veterinary sources only: the Merck Veterinary Manual, AAHA guidelines, the AVMA, the ASPCA Animal Poison Control Center, the Cornell Feline Health Center and the WSAVA guidelines. Every turn runs triage, then the medicine rule, then the answer.
| The pet guide does | The pet guide never does |
|---|---|
| Explains vaccines, heartworm, parasites, dental disease and foods that poison pets, with the source on the answer | Gives a dose, an amount or a schedule of any medicine for an animal |
| Reads the message for ten emergency patterns and sends the Parent to care before anything else is said | Suggests a medicine made for people for an animal |
| Gives both animal poison control lines for a toxin eaten or licked | Diagnoses the animal, or rules a condition out |
| Tells the Parent what to ask the veterinarian | Runs the human emergency rules on an animal's signs |
The triage rules cover a toxin eaten or licked, trouble breathing, a cat straining to urinate, bloat, seizures, collapse, trauma, overheating, a stalled labor and an eye injury. A match sends the Parent to the clinic or the nearest emergency veterinary hospital now, and nothing else is said in that turn. Each rule can only ever send a Parent toward care. The same filter runs on everything the Parent sends the clinic: a refill note or a message describing an emergency is stopped rather than delivered into a queue somebody reads tomorrow.
The human emergency rules are deliberately not run on an animal's words, because they misread an animal's signs. "Limp" is a gait in a dog. A Parent typing about a straining cat is describing a urinary obstruction that kills in hours, and a system built around human chest pain has nothing to offer.
The signoff gate
The triage rules do not go live on anyone's assurance. The rule text carries a version and a cryptographic hash of its exact text. A veterinarian's signoff is recorded against the hash of the text they reviewed, naming the veterinarian with license number and state, and a hash that differs from the running text is refused. Until a signoff exists, every pet answer says the rules are unreviewed, and in production the guide answers nothing beyond the triage instruction and "call your clinic." Any change to the rule text, down to a word, needs a new signoff.
A clinician signs off on the exact text, identified by its hash. Change one word and the signoff is gone.
The commercial argument
Pet Talisman costs the practice's Clients nothing and is included with the clinic's subscription. The practice gets a portal it did not build, under its own brand, that books visits, collects check-in forms, carries signed Treatment Plans with deposits and takes payment. The Parent gets one app instead of a paper folder and a vaccine card in a drawer.
The flywheel is the one the human side runs on. A Parent who holds their pet's complete record arrives at a new practice with it: vaccine history, laboratory values with their species ranges, the medicines the animal has had, the weight trend. That practice spends the first visit treating the animal rather than reconstructing it.
Charges that get collected
Treatment Plans signed from home with the deposit paid, payment plans that start with a charge, and refill requests that arrive structured instead of as a voicemail.
The record is theirs
It survives a move, a change of practice and a referral. Nothing in it is sold, and a Parent never pays to hold their own animal's history.
One app across every location
A group puts its whole client base on one branded app and keeps the record continuous across its general practice, its emergency hospital and its referral center.
KEELSON
Keelson, the imaging platform
The archive, the viewer, the radiology workflow and the reporting that an imaging center or a radiology group runs its day on, under its own name, priced per study.

Medical images are the heaviest and least portable data in health care. A study sits in an archive that belongs to the center that made it, is read in a viewer bought separately, is reported in a third system, and reaches the patient as a disc or a link that expires. A record that intends to hold a person's health for life cannot treat imaging as an attachment.
Keelson is the layer that holds it. It stores studies, serves them to a viewer, drives the worklist a radiologist reads from, produces the report, and shares the result with the ordering clinician and the patient. It talks the standards imaging equipment already speaks, so a scanner does not care what is behind the wall: the worklist appears on the console, the study lands in the archive, and the completion message closes the loop.
What it does
Storage that scales by study
Studies arrive from scanners and from other institutions, are stored once, and are served to any authorized reader. Prior studies for the same patient are found through the record's identity layer rather than by matching names and dates of birth, which is where cross-institution imaging usually fails.
Viewer, worklist and reporting
The radiologist opens a worklist, reads, dictates or types into a template, and signs. Priors sit beside the current study. The report goes back to the ordering clinician inside the clinic system and to the patient inside their own app, without a disc and without a link that expires.
Support, and a base to build on
Imaging AI tools attach to the workflow, and their findings enter the record as evidence with a source and a weight like any other finding. What a tool reported and what later proved true are both recorded, which is the loop that lets the record learn how far to trust that tool.
The center's own brand
An imaging center runs Keelson under its own name and colors. The imaging center configuration in the clinic system is Keelson's workflow, so scheduling, billing and the patient portal are built once and serve both.
The regulatory position, stated precisely
The archive, the sharing and the workflow need no clearance. The diagnostic viewer requires clearance before primary diagnostic reading, and so does each AI tool that detects, triages or measures. That is a funded item in this round rather than an assumption, and until it is in hand the viewer is sold for review and workflow rather than for primary reading.
Who buys it
Outpatient imaging centers, radiology reading groups and specialty clinics with their own scanners. There are about 15,000 imaging centers in the United States by Definitive Healthcare's count, and the segment for imaging archive and reporting is estimated at $6.8 billion to $7.1 billion a year worldwide. The assumed price in the model is $24,000 a year per center, which is modest against what a center pays today for an archive and a viewer bought separately.

FINARC
Finarc, the Living Financial Record
Finarc is the lifetime Evidence-Weighted Financial Record for a person and for every household, business, trust and estate they own or control. Every account, asset, policy and obligation enters once as a permanent event carrying its source and a weight for how reliable that kind of source has proven to be. Budgets, books, tax returns, loan files and estate plans are views of that record, so the work exists before anyone asks for it.

A person's money is scattered across institutions that each see one slice
A bank sees deposits and the bills that clear. A card issuer sees spending. A brokerage sees positions. A payroll provider sees wages. A tax authority sees last year. A county recorder sees the deed. An insurer sees the policy it wrote and nothing about the house it covers. Nobody holds the whole picture, and the person who could is handed a budgeting app that reads four accounts and categorizes the groceries.
The deeper problem is what each of those systems does with a figure when it arrives: every figure is believed because it arrived. A bank feed, a statement PDF, a photographed receipt, a parsed email and a figure the person remembers will disagree about one transaction, and conventional software picks one and writes it down as the truth. A net worth, a qualifying income, a business profit and an estate inventory then all read as precise, and none says which parts rest on a cleared instrument and which on somebody's recollection of a cash payment in March.
That is tolerable until a figure has to carry weight. A lender underwrites from it. A tax return is signed under penalty of perjury on it. An examiner asks where a number came from. A trustee accounts to beneficiaries with it. The question is then never what the number is. It is what stands behind it.
The same evidence weighting, pointed at money
Finarc keeps every version of a fact, weighs each by the class of source that reported it, and settles the figure from the whole set. A net worth comes with a band. A qualifying income comes with the evidence beside each line. No figure reaches an application, a return or an audit file without its provenance.
Every source a person has
Banks, cards, brokerages, retirement plans, payroll providers, tax forms, county records, insurance declarations, receipts, statement files and the person's own words, each stamped with where it came from.
Five classes, learned per source
Verified feeds and transcripts start near 0.95, statements and tax forms near 0.90, receipt reads near 0.75, what a person states at 0.60, and what Finarc infers at 0.50, shown as inferred until confirmed.
One fact, many witnesses
The settled amount is a posterior across every figure reported, including the chance that the truth is one nobody reported. Feeds drawn from one bank interface are declared as one witness and counted once.
Each entity on its own side
A person, their household, each company and each trust keeps its own ledger, keys and integrity chain. Money crosses a wall only as a matched pair that sums to zero, recorded on both sides.
The work already done
Books, tax lines, loan scenarios, insurance reviews, estate inventories and audit views are built by reading the record. A new lender format or state return is a new view, and nothing stored changes.
Checkable by an outsider
Each entity's chain is anchored outside the database every hour, signed with both a classical and a quantum-resistant signature. A proof for one event checks offline and reveals no other.
Finarc never holds or moves anyone's money. A person with authority over the entity approves anything that moves money, signs, files, cancels, disputes or sends.
Where the mathematics stops being a simulation
The trust model under both records has been measured in simulation and the result is published: a ledger accepting every value at face value held the truth inside its own stated 95 percent intervals 82 percent of the time at 25 records, 38 percent at 100 and 1 percent at 250, while learned weights held near 95 percent from about 100 records on. That is a strong result, and it is still a simulation.
Converting it to a deployment result requires independent judgments: moments when something outside the system establishes what the real figure was, so every source that reported the item is scored against it rather than against the system's own answer. Finance supplies those in quantity. A statement closes the month. A transcript arrives from a tax authority. An instrument clears. A reconciliation ties out. A lender verifies income. None came from Finarc, and each arrives monthly. A health record waits years for the equivalent, because a treatment's outcome is slow, often ambiguous, and frequently reported by the institution that chose the treatment.
Those labels matter because self-agreement is the failure mode of any system that learns which of its own sources to believe. A record that credits a source for agreeing with the figure the record itself settled on climbs toward certainty on its own echo: on a source actually worth 0.70, 400 such judgments read 0.952, against 0.773 when the weight is held to what independent labels support. Finarc is the record best placed to escape that, and the study the whole company owes is cheapest and fastest to run here.
A health record waits years for an outcome. A financial record gets one every month.
The business model has two sides, and the segment taught us why
The user pays for an adviser that knows everything about their money. The banks, lenders, insurers and advisers who serve that user pay to work from a complete, verified picture, with the user's permission.
One side alone does not carry a company. Budgeting apps see a slice of a person's money, and people will not pay much for a slice: Intuit shut Mint down in 2024 and moved its users into Credit Karma, having paid $7.1 billion for Credit Karma in 2020 to own one piece of the second side of this model, the part where financial businesses pay to reach a customer whose position is already known. Plaid was valued at $8 billion in 2026 for the bank connections alone. The published personal finance software line is $1.4 billion to $1.8 billion worldwide, which measures how little the consumer side is worth by itself.
| Side | Who pays | What they are buying |
|---|---|---|
| The member | People, households, small business owners, trustees and executors | The record with its guide, every page for every entity, and finished work: a mortgage file, a return, a set of books, an estate inventory, subscriptions ended |
| The institution | Credit unions, community banks, lenders, employers, wealth firms, bookkeepers and CPAs | Finarc under the partner's brand for its own members, and a customer served from audited figures instead of a stack of uploads, each release granted by that customer |
Prices are set with the first partners. The model assumes $10 a month from a member and about $40 a year from the financial businesses serving that member, and neither is a quoted or signed price. Finarc does not sell personal or business data, and an owner can export everything in open formats.
What is built and what the record is waiting on
Finarc runs end to end today on an invented family: a married couple, their household, a single-member design firm and a revocable trust that owns the home, with 1,553 events across five verified chains covering 18 months of bank, card, payroll, brokerage, retirement, mortgage and valuation data, plus policies, estate documents, beneficiary designations and an asset register. The problems in it were planted on purpose with known answers, so the tests check that Finarc finds each one and no others. It finds 35.
The bank and card connector speaks both of the wire shapes the United States market uses and runs against two invented institutions; the live clients refuse every call until an aggregator contract exists. Statement and downloaded-file upload works today and never depends on anyone's permission to connect. Books for every kind of business, individual and business tax returns with filing, the planning engine that finds legal tax savings, and state returns are scoped and ordered behind the current demonstration.

FINARC
Entities and audit walls
A person, their household, each of their companies and each of their trusts keeps its own record, keys and integrity chain. The entities tie together through ownership, roles and matched transactions, and they separate cleanly when one of them is examined, audited, sold or put through a lender's review.
Why one ledger per person is the wrong shape
Consumer finance software models a household. Small business accounting software models a company, one per file. A person who owns a company and a trust holds three of these products, reconciles between them by hand, and cannot answer either of the questions that matter: what is the whole position, and what does this one entity look like on its own. Both answers are needed and they pull in opposite directions. Software that merges the records cannot give the second. Software that keeps separate files cannot give the first.
The entities Finarc keeps
| Entity | What its record holds | What it produces on its own |
|---|---|---|
| Person | Personal accounts, income, retirement, credit, insurance, personal documents | Budget, net worth, personal return lines, loan files, plan |
| Household | Joint accounts and shared bills, with each member's private accounts staying private; dependents carried in a parent's record | Household budget, net worth, joint return, plan |
| Sole proprietorship, LLC, S corporation, partnership, C corporation, nonprofit | General ledger, receivables, payables, payroll, sales tax, assets, loans and owners' capital, with basis, reasonable compensation, allocations, retained earnings or restricted funds as the form requires | Books, statements, the entity return as taxed, owner schedules, payroll returns |
| Revocable and irrevocable trusts, and estates | Trust or estate property and title records; income and principal, fiduciary actions, beneficiary accounts, creditor claims and distributions | Asset schedule, successor trustee packet, accountings to beneficiaries, heirs and the court, fiduciary and estate returns |
| Clinic | A medical, dental or veterinary practice's books on the clinic chart, posted from Arca as summarized journal entries with no patient, client or employee identity | Books, statements, the entity return with the CPA package, lender and buyer views |
Entity types are defined in a registry rather than in the database structure, so a new form, a series LLC or a foreign company, is added by a definition a person approves. A change of form is an event too: a sole proprietorship becomes an LLC, an LLC elects corporate tax treatment, a revocable trust becomes irrevocable at the grantor's death, and the history before the change stays readable as it stood.
How the entities tie together
| Link | Examples | What Finarc does with it |
|---|---|---|
| Ownership | One owner holds all of an LLC; partners hold 60 and 40; a trust owns the house; a holding company owns an operating company | Builds the ownership map, rolls results up through owner schedules, and computes a consolidated position for the owner's own view only |
| Role | Owner, officer, bookkeeper, CPA, trustee, beneficiary, executor, power of attorney | Grants access to named entities for a stated scope and period, with the authority document on file |
| Matched pair | Owner draw, capital contribution, owner loan, related-party rent, management fee, trust distribution, pass-through income, wages to an owner | Records the transaction on both sides, checks that both agree to the cent, attaches a purpose and document to each, and flags any pair that is one-sided |
The consolidated view across everything a person owns is a view and only a view. It never merges the records and it is never what an outside party receives.
What a wall actually enforces
A wall in Finarc is not a filter applied when a report is printed. It is the structure of the store.
- Separate record, separate keys, separate chain. Each entity's events live in their own hash chain, and reading one entity never requires opening another. Names, institutions, account numbers, counterparties, memos and every free-text value are sealed under keys held per tenant and purpose, each ciphertext bound to its own table, row and column so a value moved elsewhere fails to open. A key per entity is scoped, which narrows a stolen key to one entity and makes erasing an entity destroy its key with every copy of its data.
- Matched pairs at every crossing. If one entity pays another, Finarc writes one event on each chain, equal and opposite, summing to zero. Each side carries its own purpose, document and random related-party token, so neither reveals the other entity's identity in an audit view. A crossing missing its other side cannot be completed.
- Commingling held at the wall. A personal purchase on the business card, a business bill paid personally, or a trustee expense paid from trust funds is held until someone repays it, reclassifies it as a draw or contribution, or documents it as reimbursable. Each hold is resolved once, and only from behind the wall.
- The record as it stood. Replaying an entity's chain to a date reproduces its books exactly as they were, including what was known on the day a return was filed.
- Clean separation when an entity leaves. A business that is sold or a trust that terminates exports whole, with its chain and documents, and the others keep their side of every past crossing.
In the demonstration family the walls carry 53 matched pairs: 18 mortgage payments of $3,420 made by the household for the trust that owns the home, 18 household contributions of $4,000 from one spouse, and 17 owner draws of $5,000 from the company. An eighteenth transfer of $5,000 with no recorded purpose sits at the wall until someone records what it was, and the company's books hold it in suspense until they do.
Who can see what
Every grant is itself an event on the entity's chain, and so is every read. An owner, spouse, trustee or executor reaches the entities their authority covers and the consolidated view across them. A bookkeeper or CPA reaches only the entities they were granted: a request for the household's plan from a role on the business returns nothing, and a test has proved that since the first release. A beneficiary sees the trust accounting the trustee shares and nothing beyond it. A caregiver works inside a parent's record under a power of attorney, with the authority document on file and its end date recorded. An examiner, auditor, lender or buyer receives a read-only view of one entity for one stated period.
That view carries opening balances rather than the transactions behind them, and shows this side of each crossing only, with the other side appearing as a random token. Every figure links to the document or feed record behind it, and the chain head is shown so the reader can verify the books as they stood. Every page the reader opens is recorded on the chain and visible to the owner. The view expires on the date the grantor sets, and its access token is returned once, at creation, with only a hash stored, so a stolen copy of the database opens no audit view. Beside it, the owner can hand over a proof for any single event that checks offline and reveals nothing else.
In the demonstration, an examination of the company's prior-year Schedule C sees 99 facts and 9 crossings from its two accounts across the granted nine months. Creating that view wrote an entry to the company's chain, and the examiner's first read wrote another.
An examination of the LLC sees the LLC. The owner's personal file is a separate grant, or it is nothing.
The structure a small business owner actually needs
An owner forms an LLC or a corporation for one reason: to put a wall between the business and the family. The wall holds only while the company is treated as separate, and commingled funds are among the first things a creditor raises when asking a court to disregard the company. The usual tooling works against the owner. Personal and business money land in the same budgeting app, the business file is reconciled by a bookkeeper who also sees the household, and the evidence that the two were kept apart is an assertion rather than a record. Trustees have the same problem with harder consequences, since state law and the Uniform Trust Code require trust property to be kept separate and the trustee to account for it.
Finarc gives the owner the position across everything and gives each outside party one entity. A lender reviewing a practice sees the clinic entity. A diligence team buying the company sees the company. Every crossing between them is a matched pair with a purpose and a document on each side, recorded when it happened, so the walls hold whether or not anyone remembers them.

FINARC
Weighing financial evidence
Every datum in Finarc carries a source class and a weight for how reliable that kind of source has proven to be at that kind of fact. The weights start from a registry and move as judgments arrive, so the record gets more accurate as it grows and can show why one figure won.
One transaction, five witnesses
A single purchase can be reported by a bank feed, a downloaded file, a statement PDF, a photographed receipt, a parsed email and the person who made it, and they will disagree: amounts by a tip or a hold, dates by the settlement lag, descriptions by whatever the merchant's processor put in the field, categories because an aggregator guessed. A conventional ledger resolves this by picking one and writing it down. Finarc keeps all of them, weighs each by the class of source that produced it, and settles the figure from the whole set, recording its confidence and every source that reported it.
The five classes and where each starts
| Class | Examples | Starts at | Good enough for |
|---|---|---|---|
| Verified | Institution feed, aggregator feed, payroll record, tax transcript, county record, custodian feed | 0.94 to 0.98 | Loan files, returns, audit views, anything a third party relies on |
| Document | Statement PDF, downloaded bank file, tax form, invoice, declarations page, signed document | 0.88 to 0.95 | Returns and applications once matched to a verified figure or accepted by the owner |
| Derived | A receipt photograph read, a parsed email receipt, an aggregator's category label, an automated valuation | 0.70 to 0.80 | Budgets, bookkeeping drafts, subscription detection, items flagged for review |
| Stated | What the person tells the guide; a decision by a person with authority sits at 0.99 | 0.60 | Planning, cash transactions, intentions, items with no other source |
| Inferred | Finarc's own estimate: a category, a recurring pattern, a probable duplicate or crossing | 0.50 | Suggestions only, shown as inferred until a person confirms them |
The ordering is not an invention. Mortgage underwriting already gives lenders representation relief when income and assets are validated from approved data sources rather than borrower-supplied paper, and auditing standards rank evidence from independent outside parties above evidence the entity produced about itself. Finarc applies the same ranking to every figure, then measures whether each source deserves its rank.
How a weight moves
Settling a fact reads the weights and writes nothing. Learning is a separate pass that records each judgment before it moves anything, and the counts persist per source and per kind of fact. A source is held by a prior of 20 comparisons centered on its own overall rate, and that rate by 40 comparisons centered on its registry starting weight, with older agreement discounted toward the present so a long clean history cannot hide a recent break. A source excellent at amounts and poor at dates carries two weights rather than one average that describes neither.
Splitting a source that finely would normally be fatal, because most pairs of source and fact type have almost no data behind them. The pull toward the source's overall rate, and from there toward the registry, means the split costs nothing where there is nothing to split.
Two properties of the settling arithmetic matter because people read the numbers it produces. A single source reports at its own weight rather than higher, since the hypothesis that the true amount is one nobody reported stands for every unreported candidate, and treating it as one inflates a lone witness, the ordinary case for a receipt with no bank line behind it. And feeds sharing an upstream count once, because a pair drawn from the same bank interface agreeing on the same wrong figure otherwise reads as 0.98 confident against a true value of 0.62.
What counts as an independent judgment about money
A weight is only worth what the judgments behind it are worth, and the judgment that matters is one where something outside the system established the real figure. Money produces those on a schedule.
A statement that closes
A monthly statement reconciles to a balance the institution computed, so every line any source reported is measured against a figure the institution stands behind. A tax transcript does the same for last year's wages, interest and contractor payments.
An instrument that clears
A check, transfer or card charge that posts and settles fixes an amount and a date beyond argument. So does a ledger agreeing with a bank to the cent, a lender verification, or a custodian statement matched to holdings.
The record agreeing with itself
With no outside reference, the only label available is whether a source agreed with the figure Finarc settled on. That confirms whatever already dominates, so it sits in its own channel at a quarter weight.
The channels are stored apart, so changing what one is worth never means replaying history. Discounting self-agreement changes how fast a weight climbs on its own echo and not where it ends, which is why the ceiling rests on the independent channel: on a source actually worth 0.70 that the record kept agreeing with, 400 judgments read 0.952 counted as truth, 0.934 at a quarter, and 0.773 held to what the independent channels support.
In the demonstration record the learned weights sit above their starting values for the sources that earned it, and the source that got something badly wrong is the one that fell.
| Source | Class | Starting | Learned | Agreed / compared |
|---|---|---|---|---|
| Institution data feed | Verified | 0.95 | 0.993 | 131 / 131 |
| Downloaded bank file | Document | 0.93 | 0.982 | 57 / 57 |
| Statement PDF | Document | 0.92 | 0.941 | 7 / 7 |
| Parsed email receipt | Derived | 0.80 | 0.929 | 36 / 36 |
| Receipt photograph read | Derived | 0.70 | 0.877 | 36 / 37 |
| Aggregator feed, either shape | Verified | 0.94 | 0.94 | Nothing to compare yet |
Across a population, the same outcomes let a new household or company start from weights learned from far more judgments than it holds, through the consent gate, only under a written research authorization, and crossing no audit wall. Weights live in the analysis layer, so a correction updates every analysis on its next run without rewriting a closed year's books.
The one correction the financial record is allowed to make
The health record holds a hard line: it never overwrites a laboratory with its own opinion. A clinician reads what the laboratory posted, as it was posted, and beside it what the evidence says about that source. No amount of cross-checking entitles a database to substitute a number the laboratory did not report.
Money is different in a specific and defensible way. A transaction has exactly one true amount, constrained by arithmetic the record can check: a pair of entries must sum to zero, a ledger must agree with a bank to the cent, a card line that posts fixes what the authorization estimated. When a receipt photograph reads $189.90 against a card line of $18.99 for the same merchant on the same day, the record is not expressing a preference. It has a figure the issuer settled and a figure an optical read produced, and one is wrong on the face of the arithmetic.
So Finarc settles $18.99, carries it into the books at a confidence in the low nineties, and lowers the weight of the receipt-reading source. The same logic reaches categories: an aggregator labeled a $6,000 deposit as income, Finarc matched it to a withdrawal from the family's own brokerage account and recorded it as a transfer, keeping $6,000 out of qualifying and taxable income.
Three limits hold that permission in place. The receipt stays in the record with its original image and reported figure intact, so anyone who thinks the card is wrong can see what it said. The correction is an event pointing at what it corrects, never an edit, so replaying the record to any date shows what was believed then. And a figure Finarc cannot resolve is not resolved: a downloaded bank file with the letter O typed in place of a zero went to a holding bay with the reason rather than into the books at a guess.
A laboratory value is a measurement of a person. A transaction amount is a fact arithmetic can check. The record is allowed to act on the second and never on the first.
Two statements about the weights hold together. Inside a record, every figure shows its sources, each source's class and the weight it carried, so the owner and anyone granted an audit view can check why one number won. The values that make the method work across many records, the tuned weights, scoring coefficients and source trust priors learned from outcomes, are trade secrets and appear in no outside document.

FINARC
What Finarc does for a household and a business
Budgets and cash forecasts, the asset register, insurance review, investment accounts, estate planning, the mortgage or business loan file, subscriptions and waste ended at one press, and a standing adviser that watches the whole position without being asked. Each service states what Finarc prepares and where a licensed professional takes over.
Four rules every service follows
- A person approves. Anything that moves money, signs, files, cancels, disputes or sends waits in an approval queue for the person with authority over that entity.
- No custody of money. Finarc never holds or moves anyone's funds. Payments and payroll run through the entity's own bank or a licensed partner.
- Licensed work stays licensed. Where the law reserves work for a licensed party, Finarc prepares the file and hands it over, or holds the license itself first.
- Every page, every entity. One bar switches the budget, books, loans, taxes, insurance, investments, assets, estate and plan pages between a person, a household, a business and a trust.
Budget and cash flow
Where every dollar went, a plan by category, and a 90-day forecast built from recurring charges, expected deposits and recent spending. The forecast is the part conventional budgeting does not attempt, because it requires knowing which charges recur, on what cycle, and against which deposits. In the demonstration joint checking dips to $377.12 on October 1, $122.88 under the family's own cushion, because the $3,420 mortgage posts one day before payroll, and the drafted fix is a date change. The same engine finds money in the wrong place: $69,506.47 earning 0.01 percent while $4,371.76 on a card costs 24.99 percent, about $1,092.50 a year for nothing.
Subscriptions and waste, cancelled at one press
Finarc finds every recurring charge across personal and business accounts, including annual and app-store billing, from transactions and parsed email receipts. For each one the owner wants gone it drafts the cancellation, and on approval acts as authorized agent through the merchant's own path and records the confirmation on the entity's chain. Where the card issuer supports merchant-level blocking or single-merchant virtual cards, it sets one up so a cancelled service cannot keep charging, and a charge that posts anyway becomes a drafted dispute with the confirmation attached.
The same detectors catch what cancellation does not: a $1 trial that became a $119.88 annual charge, a streaming price raised from $15.49 to $17.99 without notice, a duplicate $86.40 charge on one day, an overdraft fee, and a $340 printer repair paid while a five-year service plan was in force. The five standing subscriptions run about $4,368 a year.
The asset register
Every vehicle, piece of equipment and valuable, with where it is, its serial number, what it cost, how it depreciates on the federal and state returns, what it would cost to replace, which warranty covers it, and who holds a lien on it. Depreciation follows the published tables with the bonus, mid-quarter and passenger-vehicle rules, and a sale moves to the right form rather than being booked as revenue.
This is the register an insurer, a lender or a buyer asks for and almost nobody keeps. In the demonstration it finds a lien still on file 18 months after the loan was paid, a laptop sale booked as revenue that belongs on a disposition form, a ring appraised at $14,500 against a policy that pays $1,500 for jewelry theft, and $24,899 of business equipment with no property coverage.
Insurance review and quotes
Finarc keeps every policy the family holds, whichever entity owns it, and matches each premium to the bank payment that pays it. It then checks coverage against what the family owns and earns: an umbrella against net worth, disability against self-employment income, life against ten years of income plus the mortgage, a home titled to a trust but insured in the owners' names, a dwelling limit below a rebuild estimate. In the demonstration that produces eight gaps, including no umbrella against about $1,046,535 of exposed net worth, a home insured $187,500 to $402,500 below rebuild cost, and no disability coverage on $107,220 of business income.
Nothing leaves until the person approves a release to one licensed producer they choose, recorded on the entity's chain. Quotes come back with the carrier's financial strength rating, the producer's license and compensation, lined up like for like: the lowest comparable quote is $5,405 a year, and a cheaper $4,730 quote is marked not comparable because it insures the dwelling for $150,000 less. Finarc earns no commission.
Investment accounts
Holdings arrive from custodians and plan record keepers and are reconciled to each statement balance. Finarc shows allocation across the family, fund costs in dollars a year against a broad index fund in the same asset class, gains and losses by tax treatment, losses available to offset gains with the wash-sale rule stated beside them, concentration in one company, and contribution room under the published limits. Security selection stays with the person or a registered adviser. What Finarc contributes is the arithmetic nobody does: $228,661 across three accounts costing about $359 a year in fund fees, one stock at 30 percent of a brokerage account, a $1,450 loss available against gains, and a retirement plan through the business worth about $4,384 a year in tax that the owner has not opened.
Estate planning
Finarc keeps the inventory an estate attorney asks for and charges to assemble: which documents exist, when each was signed, who the fiduciaries are, and where each original is kept. It shows how every asset passes, by the trust, by designation, to a joint survivor or through probate, checks every designation for missing contingents and minors named directly, and adds up what would pass through probate at the second death against the state's small estate limit.
In the demonstration about $374,009 would go through probate at the second death, above California's $208,850 limit, with statutory fees near $20,960, because joint accounts carry no payable-on-death beneficiary and the LLC is owned outside the trust. It also finds a term policy naming a minor child directly as contingent beneficiary, which usually forces a court-appointed guardian of the estate, and a missing durable power of attorney. Finarc drafts the request for the attorney and prepares a successor trustee packet released through a scoped audit view. It does not draft wills or trusts.
Mortgage and loan applications served from audited figures
The loan page builds the household's mortgage file from the records. For each purchase price it computes principal and interest, the full payment with taxes and insurance, front and back debt-to-income, cash to close, and the liquid assets available, with every line tied to the evidence behind it, and it lists what a lender will ask for that the record lacks, such as two full years of self-employment returns where it holds 17 months of books.
| Price | Loan | Payment with taxes and insurance | Front DTI | Back DTI | Cash to close |
|---|---|---|---|---|---|
| $650,000 | $520,000 | $3,989.96 | 20.8% | 41.2% | $149,500 |
| $825,000 | $660,000 | $5,023.79 | 26.2% | 46.6% | $189,750 |
| $1,000,000 | $800,000 | $6,057.63 | 31.6% | 52.0% | $230,000 |
The figure that shows what the record is for: qualifying income comes out at $19,180.96 a month from verified salary and a reconciled business ledger, against the $22,000 the family stated, and the file uses the records. Liquid assets of $118,996.82, with the brokerage counted at 70 percent, cover none of the three scenarios, and the page says that too. A loan file that overstates the borrower is worth nothing to either side. Other entities get their own loan views: a person borrowing alone, a business's debt capacity after owner draws, a trust's equity line capacity. Finarc prepares the file and the person sends it, since taking applications or accepting lender fees brings originator licensing and referral-fee limits.
The adviser that does not sleep
Detectors run on every entity whenever new data arrives, and each finding carries its evidence, the dollars at stake and a drafted action. That is the difference between software a person visits and an adviser who notices. The record produces 35 findings across five entities in four bands: money being lost now, items held at an audit wall, fixes to make, and items recorded so a choice gets made on purpose. A human adviser who found all 35 in one engagement would be a good one, and would charge for the hours.
The file uses the records. Where the family's stated income was higher than the evidence supports, it says so.

FINARC
Books, returns and applications
Everything a person, an accountant, a lender or an examiner sees in Finarc is a view built from the events: books by account and period, tax lines tied to the year they belong to, loan scenarios, and the chain that proves nothing was altered. Thirty-three report pages run across the five entities of the demonstration record.
Views out, events in
Each fact is written once as an event, and views are shaped for their readers: the budget by category and date, the books by account and period, the loan file in the industry's own data format, the return in the lines the filing system expects. Keeping the event store separate from the views means a new report, a new lender format or a new state return is a new view built by reading the record, and nothing stored changes to produce it.
A filed return is reproduced exactly by replaying the record to the filing date, and an amended return is a new event pointing at the original. That lets an accountant answer the hardest question in the practice: what did the books say on the day I signed, and what has changed since.
Bookkeeping, for whatever the business is
Every entity gets a double-entry ledger with debits equal to credits, for the current and prior year. Personal and household books carry investments and property at market value. Business books hold commingled items in suspense until someone resolves them, so the ledger never quietly absorbs a personal charge as a business expense.
In the demonstration, the design firm's books through September show revenue of $111,941.74 against contractors, rent, equipment and depreciation, with net income of $79,112.59, $160.16 held for commingling and $5,000.00 in suspense for an unrecorded crossing. Debits and credits each total $417,104.39. The trust's books carry the home at $1,182,800 and the mortgage at $579,534, with $61,560 of payments made on the trust's behalf by the household recorded as equity paid by a related entity rather than as rent or a gift.
The ledger built today runs a single-member service business, a household and a trust. Industry breadth is the next body of work and it is specified: a chart of accounts pack per industry in the pattern of the clinic chart, from professional services and retail with inventory to construction with job costing, rentals, farms, trucking and nonprofits with restricted funds, with classes and locations, sales tax by jurisdiction, accrual and cash books side by side, and payroll for any employer rather than only for clinics.
Tax returns, prepared and optimized within the law
The federal projection is built and runs year round rather than in April. Every figure that depends on the year, brackets, standard deduction, wage base, contribution limits, the estate exclusion, is looked up for the tax year being projected rather than for today, and only years the revenue service and the social security administration have published are on file. A projection for a year not yet published uses the latest year on file and says so; nothing is extrapolated.
| Line | Amount | Where it comes from |
|---|---|---|
| Wages, annualized | $111,772.89 | Payroll provider records, 19 pay periods, less deferrals |
| Business profit, annualized | $107,220.08 | Business ledger, $79,112.59 to date, less depreciation from the asset register |
| Adjusted gross income | $211,418.11 | Computed, after half of self-employment tax |
| Deductions | $52,129.04 | The published joint standard deduction, plus the business income deduction |
| Total federal tax | $39,617.31 | Income tax of $24,467.60 on $159,289.07, plus self-employment tax of $15,149.71 |
| Paid to date, against $26,741.68 required | $24,628.00 | Payroll and bank records: a $2,113.68 shortfall, caught in September rather than April |
The projection annualizes 273 days of records by day of year, and uses the standard deduction until the county property tax bill is on record, with the trust's tax page stating the exact point at which itemizing wins: property tax above $7,036.
Lowering what a taxpayer legally owes is a separate engine from computing what they owe, and it is specified to run across the whole record year round: filing status, itemizing, retirement and health savings contributions, the timing of income and deductions, depreciation elections, entity choice and the corporate tax election, reasonable compensation, charitable strategies, loss harvesting, credits and state residency. Every recommendation shows the rule it rests on, the dollars it saves, what it costs or risks, and any disclosure it requires. A position needing disclosure is labeled and never taken by default, and no position the law does not support is offered at all.
Three ways a return gets finished
| Route | How it works, and what it waits on |
|---|---|
| The accountant reviews and files | A scoped, logged, expiring grant into the entity, every figure one step from its source, proposed adjustments, approval by name, and the hand-off in the formats professional tax software imports. Built for clinic entities today |
| Finarc files for the taxpayer | The taxpayer reviews and signs electronically, Finarc transmits through the federal and state filing systems and resends what comes back rejected. Scoped: it requires authorized provider status as software developer, transmitter and return originator, with suitability checks and assurance testing each season, plus each state's approval |
| A preparer prepares and signs | For people who want a human preparer. Scoped, and not a software question: each preparer needs an identification number and, in California, state registration unless they are a CPA, attorney or enrolled agent |
Individual returns are specified to cover the main form with every schedule behind it, including capital gains with cost basis from the brokerage record, rentals and pass-through income, retirement distributions, credits and health savings accounts, plus every state's returns. The difference from consumer tax software is the starting point: consumer software collects a year of documents in a few weeks from a taxpayer who is guessing what matters, while Finarc keeps the books all year, so filing confirms what is already there.
A balance due is paid from the taxpayer's own bank account through the filing payment channel, the government's direct payment service, or by card through the approved processors. Refunds go from the government straight to the taxpayer's accounts by direct deposit, including split refunds; the government pays, and Finarc routes and tracks it. Taking Finarc's own fee out of a refund is a bank product needing a bank partner and its own disclosures, and it is a later choice rather than a default.
The clinic books that Arca feeds
A medical, dental or veterinary practice's money splits at the line where patient identity stops. Patient and client accounts, charges, invoices, claims, remittances, deposits and statements stay inside Arca. The practice's books are a Finarc entity of their own, and Arca posts to it only summarized journal entries: charges by revenue class, payments by type, deposits, discounts, write-offs, refunds, sales tax collected and the cost of goods relieved from inventory, as one netted entry per account each day.
Every entry must balance and every account must be in the clinic chart. An entry whose references carry anything naming a patient or an employee, a chart number, a record token, a birth date, a staff name, is refused at any depth, and no diagnosis or procedure codes reach Finarc, which keeps the financial record outside federal health privacy scope entirely.
The entity produces a full set of books: profit and loss by month, location and provider; a balance sheet; direct cash flow; payables; fixed assets; bank reconciliation; a month-end close that refuses to run while a bank is unreconciled; and the tax package for the clinic's CPA. Arca and Finarc carry the same chart of accounts file, byte for byte, with Finarc recording its digest on the chain so neither side can drift. The owner's personal entity, a management company and a real estate entity leasing the building each sit behind their own wall, and rent, fees and draws cross only as matched pairs.
A clinic's books reach Finarc as balanced journal entries with no patient in them. That is why the financial record stays outside health privacy scope.
What the chain proves about a report
Every report can be checked against its entity's chain, which is anchored outside the database every hour under both a classical and a quantum-resistant signature. All five chains of the demonstration verify as loaded. A lender, auditor or court can also receive a proof for one event that checks offline and reveals no other. For an accountant signing a return or a trustee accounting to a beneficiary, that is the difference between asserting that the books were not altered and demonstrating it.

FINARC
Connecting to the institutions
Accounts, balances and transactions arrive from banks, card issuers, brokerages, retirement plans, payroll providers and insurers, each under a consent one person grants for one entity at one institution. One adapter speaks both of the wire shapes the United States market uses, so a second provider is a configuration change rather than a rebuild.
What the connection is for
A budgeting app connects to accounts so it can draw a pie chart. Finarc connects so that a figure in a loan file, a return or a trust accounting rests on something an institution reported rather than on something a person typed. The connection is the top of the evidence hierarchy, and everything downstream depends on having it: the confidence band on a net worth, the qualifying income a lender can rely on, the reconciliation that closes a month. It is also what produces the independent judgments the trust model learns from. A posted line that contradicts a receipt read, a statement that closes against a month of feed lines, a pending authorization that posts at a different amount: each is a label on a source's reliability that did not come from the system's own answer.
What arrives, and from where
| Source | How it connects | Class and starting weight |
|---|---|---|
| Banks and card issuers, through a data aggregator | The person signs in at the institution's own page and the aggregator returns accounts, balances and transactions. Both market shapes are implemented | Verified, 0.94 to 0.95 |
| Direct download | The standard download formats a bank still offers, or the file the person downloads themselves | Document, 0.93 |
| Brokerage and retirement custodians | Holdings, cost basis and transactions, or a custodian feed under an adviser agreement | Verified, 0.97 |
| Payroll providers | Employment and income records through consumer-permissioned access, or an uploaded wage statement | Verified, 0.97 |
| Tax authority transcripts | The taxpayer downloads from their own account, or a lender uses the income verification service | Verified, 0.98 |
| Insurance carriers | Declarations pages and benefits statements by upload; quotes return from a licensed producer | Document, 0.88 to 0.90 |
| Statements and receipts by upload | Statement PDFs, tax forms, paper and emailed receipts read with text recognition | Document 0.92, derived 0.70 to 0.80 |
| What the person says | A conversation with the guide: cash spending, debts owed to them, intentions | Stated, 0.60 |
Customer-approved recurring access is what separates this from an annual document-gathering exercise. With it, a reconciliation closes every month instead of every April, a cash forecast has something to forecast from, a subscription price rise is caught in the month it happens, and the weights learn monthly. Without it, Finarc is a very good filing cabinet.
How a consent works
- One consent, one entity, one institution. A connection is created by a person, never by a service, who holds an authority role on that entity or on an entity that owns it. A bookkeeper, CPA or beneficiary may read a record and cannot open its accounts.
- Named scopes, a year at most. The consent names what it covers, accounts and optionally balances and transactions, and lasts no longer than 365 days.
- Walls at the link. Only the entity's own accounts are linked, matched by account number through a keyed index that never decrypts anything. An account at the same institution belonging to another entity is listed with that reason and its lines are never stored, so a household's link cannot pull the company's card into the household's record.
- On the chain. The grant, every sync and every revocation are events on the entity's own chain, so the owner can see what was connected, when it last ran and what it can read.
- Revocation that means something. Revoking deletes the access token at once and tells the aggregator, and consent withdrawn at the institution is detected on the next sync. Nothing is imported after either, and data already received stays, because it is the owner's record.
Access tokens, cursors, institution and account identifiers, masks and names are all sealed at rest under keys held per tenant and purpose, each bound to the connection it belongs to, and the token is cleared when consent ends. What stays readable is structural: identifiers, adapter, status, scope codes, and the consent and sync times shown to the owner as a timeline.
Posted, pending and the evidence each becomes
A posted line becomes an observation event carrying the aggregator as its source, and the reconciler groups it with the same line from any other source exactly as it does for a statement or a receipt. A line reported by both the aggregator and the statement is one fact with two pieces of evidence, and feeds that read a single bank interface are declared as one source family and counted once.
A pending line is a separate kind of event, shown to the owner and kept out of the books, with the posted line that replaces it pointing back at it. That distinction is why a card feed does not corrupt a cash forecast: a fuel authorization of $100.00 that posts at the pumped amount, and a hotel hold released after three days without posting, are both ordinary and both wrong to book. Statement upload sits beside the connector permanently rather than as a stopgap, since it works at institutions that refuse or throttle programmatic access, and the locator for each file is a hash of its content, so the same file imported twice adds nothing.
The partnership route
Finarc does not contract with thousands of banks. It contracts with a data aggregator, and the aggregator holds the agreements with the institutions. The candidates differ in ways that matter: one is a network owned by a group of large banks and a clearing house that passes data through without storing it, and the others carry broader coverage of smaller institutions. The choice turns on coverage of the institutions the first members hold, price per connection, and whether the provider will sign the data-protection terms the suite requires.
Supporting both wire shapes from one adapter answers a commercial risk rather than an engineering preference. The federal rule that would give consumers an enforceable right to their own financial data was finalized in October 2024, challenged in court, disowned by the agency that wrote it, enjoined pending reconsideration, and sent back for a revised proposal in August 2026 whose content is not public. The April 2026 compliance date passed with no obligation in force, and large institutions have begun charging aggregators for access, with the first paid agreement reached in September 2025.
So no rule compels any bank to hand Finarc data, and access runs on contracts. The plan assumes per-connection fees, assumes some institutions refuse or throttle, keeps two wire shapes live so a second provider is a configuration change, and keeps statement upload as the path that always works.
No federal rule compels a bank to hand us data. Access runs on contracts, which is why there are two wire shapes and an upload path that never needs anyone's permission.
What is real today
The connector is built: both wire shapes normalized to one sign convention before anything is stored, consent, sync, revocation, walls at the link, sealed tokens, and the tests that hold all of it. It runs against two invented institutions whose records continue the synthetic household forward as the date moves, so they behave the way real ones do: pay every other Friday, monthly bills, card spending and interest, a line pending for two days and then posted, fuel authorized at $100.00 and posted at the pumped amount, a hotel hold released without posting, and a card issuer that withdraws consent three weeks after the link so withdrawal at the institution is exercised on every run.
No real bank, card issuer or aggregator has been contacted, and no real financial data has entered Finarc. In live mode the clients refuse every call and the interface answers with a not-implemented status, so nothing is sent anywhere.
| Before a live connection | State |
|---|---|
| An aggregator contract and credentials | None exists. The live clients refuse every call |
| Aggregator callbacks, to shorten the delay between a posting and its arrival | Scoped. Finarc polls on a schedule today |
| Matching after a re-link, since a real provider issues new identifiers | Scoped: lines already imported matched by date, amount and description |
| The rules that apply before real data: a written security program under the federal safeguards rule, a risk assessment, an outside penetration test | Scoped, and gated ahead of the first live member |
Custodian, payroll and tax-transcript connections follow the same gate. A first white-label partner requires a completed Type II examination, which is what a credit union asks for before putting its name on software that touches its members' accounts.

MY HELM
My Helm
My Helm is the family's app for the financial record, the money counterpart to My Talisman. A household talks to it about any account, bill, goal or tax season, and it answers from Finarc's verified figures, prepares the work, and hands anything the law reserves to licensed hands. Its mark is a ship's wheel, because the family steers and the record holds the course.
Why the record needs an app of its own
Finarc is the record and the engine. It holds the events, weighs the evidence, keeps the walls between a person and their companies, and produces books, returns, loan files and audit views. None of that is what a family wants to look at on a Tuesday evening.
The health side of the suite settled this question already. Holarc holds a person's health for life and patients never see it; they meet it through My Talisman, which carries the family and does the work of staying healthy with them. The financial side needs the same split for the same reason: a record built to satisfy an examiner and a record a household will actually open are different products with different jobs, and collapsing them produces software that does neither well.
So My Helm is the hand on the wheel. One conversation covers the household and the business, and every number it reports came from the record with its evidence attached, which is the part a budgeting app cannot do and the part that makes the answers worth acting on.
What a household sees
The home screen earns its place by reporting what changed since the family last looked: position across everything, what needs a decision, what is about to go wrong, and what has already been prepared and is waiting for approval.




The guide, and one press
The guide is the way into all of it, by voice or by typing, and it answers from the record rather than from a general model's recollection of personal finance. It states which entity it is working in before it answers, because the same question has a different answer for the household and the company, and it asks before recording anything that crosses a wall between them. What it will not do is as defined as what it will: it drafts, a person with authority approves, and it never holds or moves money, advises on specific securities, drafts a will, sells insurance or files a return on its own authority. Voice is never mandatory.
The actions are the product. A household that can see its position and do nothing about it has a dashboard. My Helm's screens each end in something a person can approve.
End a subscription
The cancellation is drafted, sent through the merchant's own path on approval, the confirmation recorded on the entity's chain, and a merchant block or single-use card set where the issuer supports one.
Claw back a charge
A duplicate, a fee, a charge after cancellation or a repair paid under a warranty becomes a drafted dispute or refund request with the evidence attached.
Fix the tax shortfall
An estimated payment now, or higher withholding for the rest of the year, which counts as paid evenly across the year and covers the penalty as well.
Send a complete loan file
Income, assets and debts tied to their evidence, debt-to-income at several prices, cash to close, and what the lender will ask for that the record does not hold.
Get like-for-like quotes
A release to one licensed producer the family picks, recorded on the chain, and quotes lined up on equal terms with a cheaper one that insures less marked as not comparable.
Open a door and close it
A scoped, logged, expiring view for a CPA, a lender, an examiner or a successor trustee, covering one entity for one period, with every page they open recorded.
The family and the entities, in one app
A household is rarely one legal person. My Helm carries each member, the household itself, each company and each trust, and switches between them with one control while reading only the entity selected. Spouses share the accounts they choose to share and keep the rest private. Dependents are carried in a parent's record. An adult child with a power of attorney manages a parent's finances inside the parent's record, with the authority document on file and its end date recorded.
The consolidated position across everything a family owns is a view for the family and only for them. It never merges the records, and it is never what an outside party receives. A crossing between two entities, an owner draw, a household contribution, a mortgage the household pays for the trust that owns the home, appears on both sides as a matched pair with its own purpose and document, and anything unexplained sits at the wall until a person records what it was.
For a small business owner this is the whole point. The same app shows the family's position and the company's books, and the wall between them is the structure of the record rather than a setting. The liability shield an LLC exists to provide survives only while the company is treated as separate, and the evidence of that is now a record made at the time rather than a reconstruction made under pressure.
The route where the institution pays
My Helm reaches households two ways. A family can subscribe directly. Or a credit union, community bank, small business lender, employer or wealth firm offers it to its own customers under its own brand, and pays for it.
| Direct | Offered by an institution | |
|---|---|---|
| Who pays | The household, with per-entity pricing for businesses | The institution, per partner license |
| Whose brand | My Helm | The institution's, from one code base set by a brand pack |
| Whose customer | The household's own relationship with Finarc | The institution's, and it stays the institution's |
| What the household gets | The full record and every page for every entity | The same product, at no cost to them |
The institution gets members who arrive with audited figures, loan files that arrive complete, and a guide that works across every institution its member uses. That route is how the suite reaches households without a consumer marketing budget: each partner brings its whole customer base on the day it signs, the same distribution argument that puts Arca in clinics rather than selling a health app one download at a time. It also fits what a credit union needs, since it cannot build this, its members are using somebody's budgeting app already, and a member whose position is verified is one it can lend to faster.
The relationship stays the institution's. Finarc does not sell personal or business data, takes no insurance commission, and takes no lender referral fee unless it is lawfully structured. An owner can export everything in open formats at any time. What the partner buys is the software and the access its own customers grant it.
A budgeting app shows a household what it spent. My Helm hands it the work, already done, waiting for one approval.

FINARC
The Finarc market
People, households and small businesses subscribe. Trustees and executors buy administration packages. Credit unions, community banks, lenders, employers and wealth firms license the record under their own brand and pay to serve their own customers from verified figures. The consumer line alone is small, which is the reason the model has two sides.
The market Finarc sells into
Financial records, aggregation and small business accounting, worldwide.
Sizing, from published estimates
The segment Finarc sells into directly, financial records, account aggregation and small business accounting, is estimated at $16.5 billion to $16.9 billion a year worldwide, and is the financial share of the suite's $107 billion to $128 billion core total. The tiles below carry each line's high estimate.
These are pooled from published estimates by Fortune Business Insights, Grand View Research, Precedence Research, Mordor Intelligence, MarketsandMarkets, Global Market Insights, IMARC and others rather than from one commissioned study. Research firms define these segments differently, so each line is a range between the low and high estimate, several overlap, and adding them together would count some spending twice.
Why the consumer-only line is the small one
The $1.4 billion to $1.8 billion for personal finance software measures how little a household will pay for a slice of its own money, and the reason is structural rather than a failure of execution. A budgeting app reads four accounts and hands the person a chart. It cannot prepare a return, because it does not hold the books. It cannot assemble a loan file, because it cannot stand behind a figure. It cannot keep a trust accounting or wall a company off from a household, because it models one household and nothing else.
Intuit proved the point with the clearest data anyone has. Mint was the best-known product in the category and Intuit shut it down in 2024, moving its users into Credit Karma, which the same company had bought for $7.1 billion in 2020: a free consumer product whose revenue comes from the lenders and card issuers that pay to reach a customer whose position is already known. Plaid was valued at $8 billion in 2026 for the bank connections alone, without the record on top of them or any consumer relationship. Those three numbers locate the value in this segment, and it is not in the chart.
The two revenue sides
| Line | Buyer | What they get | Structure |
|---|---|---|---|
| Consumer subscription | Individuals and households | The record with its guide and every page for every entity they can see | Individual and household tiers |
| Small business subscription | Owners of sole proprietorships, LLCs, S corporations, partnerships | Books, tax and payroll records per entity, each behind its own wall | Per entity, discounted for owners of several |
| Trust and estate administration | Trustees and executors | Trust accounting, fiduciary returns, successor packet, estate inventory | Per package |
| Single-purchase packages | Anyone facing one heavy lift | A completed mortgage file, a completed return, a sale or diligence package | Per package |
| White-label license | Credit unions, community banks, lenders, employers, wealth firms | Finarc and My Helm under the partner's brand for its own customers | Per partner license |
| Professional seats | Bookkeepers, CPAs, advisers, trust officers | Scoped access to the entities they are granted | Per seat |
| Research licensing | Researchers | De-identified answers, only under a written research authorization, through the consent gate the health side uses | Per license |
The modeled assumptions are $10 a month from a member and about $40 a year from the financial businesses that serve that member, and neither is a quoted or signed price. Finarc earns no insurance commission, takes no lender referral fee unless lawfully structured, and does not sell personal or business data; the value is the software and the access a customer grants.
The member side is what gets a household to sign up and what produces the independent judgments the record learns from. The institution side is what makes the business large. Both need the same thing to exist, a complete verified position per entity, and that is the asset neither a budgeting app nor an aggregator holds.
Who buys, and why they buy
People and households get a mortgage file, a tax return, a budget, a plan, an insurance review and an estate inventory out of records they already keep, with subscriptions and waste ended on approval. The arithmetic is easy to check: in the demonstration record one finding alone, a retirement plan the business never opened, is worth about $4,384 a year in federal income tax, and the five standing subscriptions run about $4,368 a year.
Small business owners keep each company walled off from their personal money, which protects the liability shield the company exists to provide, and hand a lender or an examiner one entity for one period. Trustees and executors keep trust or estate property separate and produce accountings from a record that can prove it was not altered. Institutions get members served across every bank those members use, loan files that arrive complete, and a product they could not build.
The comparables, and where each one stops
The closest incumbents are each built around one product they sell. None keeps one lifetime record with provenance and confidence on every figure across a person and every entity they own, with walls between those entities enforced by the record itself.
| Who | Where Finarc differs |
|---|---|
| Intuit: TurboTax, Credit Karma, QuickBooks | Separate files for separate jobs, with Mint's users moved into Credit Karma when Mint shut down in 2024. Finarc is one record across the person and every entity, and the books feed the return rather than being collected for it |
| Plaid, MX, Finicity, Akoya | Suppliers rather than competitors. Finarc speaks both wire shapes, and the record, the weighting and the walls sit above them |
| Rocket Money with Rocket Mortgage | Rocket's budgeting app feeds its own lending. Finarc prepares the file for whichever lender the person chooses |
| Lender-owned assistants | Licensed to a mortgage company to steer people toward that lender's loans. Finarc has no product at the other end of the advice |
| Xero | One company per file, with bank feeds treated as exact and no second witness to any figure |
| Monarch, Copilot, YNAB | A household at most, with categories a person edits and no weight, provenance or confidence anywhere |
| Empower | Investment tracking alone, where Finarc covers investments as one page among budget, books, loans, taxes, insurance and estate |
One opening is worth stating plainly: the government ended its own free filing tool after the 2025 season, so no government-run free filing option remains in 2026. That leaves room for returns built automatically from records their owners already keep.
What none of them combine is one lifetime record per person and per entity, every event kept, chained and replayable to any date; a learned weight and its provenance on every figure, with a confidence band on every total; audit walls enforced by the store; no product of its own to steer a person toward; and one weighting method shared with the health record. A late entrant can copy the software and still starts with no accumulated evidence for the weights to learn from, and what a conventional ledger rounded away at entry cannot be recovered later at any price.
Mint shut down in 2024. Credit Karma sold for $7.1 billion. The difference between those two numbers is the whole argument for a two-sided model.
The first partners, and the gates in front of them
The first institutional partners are credit unions and community banks, for a reason that has nothing to do with sentiment. They hold the member relationships, they cannot build this, their members are already using somebody's budgeting app, and a member whose position is verified is one they can lend to faster. Small business lenders and employers follow, then wealth firms. Professional seats sell alongside all of them, because a CPA granted scoped access into a client's entity is the cheapest distribution there is.
| Risk to the plan | Response |
|---|---|
| Bank data access rules are unsettled: the federal rule is enjoined and being rewritten, and large banks have begun charging for access | Contracts with an aggregator rather than with banks, a plan for per-connection fees, one adapter with two wire shapes so a second provider is a configuration change, and statement upload as the path that always works |
| Licensed work: tax filing, loans, insurance, securities advice, law | Licensed partners until Finarc holds the license itself, each regulated service behind its own gate |
| Partners require assurance before putting their name on it | A completed Type II examination before the first white-label partner goes live, plus tamper evidence a partner, auditor or court can check offline |
| The name | A fee-only wealth manager uses the FinArc name in financial services, so a different public name may be needed before launch |

NERD
NERD, the evidence engine
NERD is the Numerical Evidence and Research Discipline: the engine that scores claims, weighs studies, runs research inside the walls and returns answers. It tells a researcher what the data can support, what it would take to support more, and it keeps every result a lab produces, including the ones that failed. It is also the engine that produced this company's own products, which makes it the deepest asset Quantavera holds.

What it is
A researcher brings NERD a question, a dataset, a candidate design or a half-formed idea. NERD reads what the published record already says and scores it source by source. It helps decide what to measure and how many subjects it will take. It chooses and runs the analysis, reports an effect size with an interval and its assumption checks, scores how close the study has come to knowledge, and names the cheapest next step. Then it remembers all of it: the result, the design, the person who judged it, and whether the judgment turned out to be right.
Every verdict, every claim, every gate and every export is decided by a named person, and the software has no path to decide any of them on its own.
NERD advises and remembers. People decide.
How research decides what is true, and what that costs
Research decides what enters the record with a single threshold. A result lands on one side of a p-value of 0.05 and becomes publishable, or lands on the other side and mostly disappears. Three things follow, and all three are expensive.
The threshold answers a question nobody is asking. A p-value reports how surprising the data would be if the effect were exactly zero. A researcher deciding whether to spend three months and a budget on an idea needs the probability that the effect is large enough to matter, the probability it is too small to ever matter, and whether the next affordable test could change the answer. The threshold supplies none of them.
Failures are never recorded. A study that found nothing tells the next person where not to go, and the system throws it away. Every lab pays repeatedly for the same dead ends, and the published record it reads is selected for positive results.
There is no way to retrace a conclusion to the claims underneath it. A claim that seems to hold becomes the ground for later papers, and those become the ground for more. When something far up the chain turns out to be wrong, the cause is usually a claim two or three layers down that was incomplete rather than false: a dependency nobody declared, a condition nobody stated. The original paper cannot be changed and keeps being cited.
What NERD does instead
NERD replaces the single threshold with four instruments that answer the questions a working researcher and a working company actually have.
The NERD Score
Every registered study carries a score from 0 to 100 across six dimensions, with design scored by the class of study run, and a ranked list of the cheapest next steps that would raise it. The score states what the design can support and what it cannot.
The Verdict
Pursue, Probe, Pivot or Park, from the probability a path exists, the probability it is a dead end, and whether the next affordable test would change the answer. A person judges every verdict, and every verdict is later scored against what happened.
The claim web
Claims are held as a dependency web, so a failed test becomes evidence about the premises underneath it and a notice to every other claim resting on the same ground. Nothing changes a claim's confidence without a person accepting it.
Retained failure
Nulls are kept, flagged and credited. A stop carries its numbers and its lesson. A lab's memory of what did not work becomes an asset it can search instead of a cost it pays twice.
Four properties run through all of it: provenance on every number, confidence calibrated against outcomes, failure retained, and validation by someone outside the authorship. Three judgments stay with people: what a result means, what gets built, and what passes review.
The two businesses NERD supports
NERD earns in two ways, and the second one is the reason it sits at the center of the company.
Research answers served from the records
Holarc and Finarc hold lifetime records in which every fact carries its source, its original document and a learned weight. A research sponsor, a device company, a drug developer, a university or a health system, buys a defined question run by NERD inside our walls on consented, de-identified data. The sponsor receives scored answers, negative results included, with the NERD Score and the Verdict attached and a plain statement of what the design can support. The sponsor never receives records, identities or a dataset, and an answer is released only after a check that no small group in it can be singled out. Patient and customer data is not sold.
A broker sells a file. NERD sells a question answered under a discipline, with the nulls the sponsor's own program might have buried, and the records never leave.
| Term | Research project license |
|---|---|
| Buyers | Device companies, drug developers, universities, health systems |
| What they get | A defined question run inside our walls on consented, de-identified data; scored answers with negative results included |
| What they never get | Records, identities or a dataset |
| Pricing basis | Per project, by cohort size and scope, renewable as the data grows |
| Release rule | An answer is released only after a check that no small group can be singled out |
NERD also serves the records themselves. It learns from outcomes how often each kind of source turns out to be right, a clinic blood pressure, a wristband reading, a typed entry, a bank feed, a scanned statement, and returns those trust weights and scored claims to Holarc and Finarc. Each record already learns its own sources for its own kinds of fact. NERD adds the level across records, which no single clinic or bank can reach alone.
The engine that produces this company's products
The second business has no invoice attached, and it carries more value. The methods NERD encodes are the methods the founders used to build the suite: every design claim carrying a confidence score and its provenance, failed approaches staying on the record so they are never repeated, independent reviewers checking the work, a person making each call and the call scored later, and the claim web tracing which decisions rest on which claims. Eight products across two regulated industries came together as one coherent system under that discipline, and the same discipline applies to every product that follows.
A competitor can read a description of the method and will still be running a conventional development organization at conventional cost. NERD is how the next product gets built, and Quantavera owns it.
What is real today
The engine exists and runs. Its statistical core reproduces published reference results on four public datasets: the Palmer penguins, the Titanic passenger list, the NCCTG lung cancer survival data and the thirteen BCG vaccine trials, and any change to a statistical method has to keep those four checks passing. Scoring, the Verdict under three priors, the claim web, retained failure, the classification layer and the AI team rules are built, sealed at rest, and recorded on a hash-chained log anchored outside the database.

NERD
The NERD Score
The NERD Score is the Epistemic Confidence Score: a number from 0 to 100 that measures how far a study is from knowledge, states plainly what the study's design can support, and lists the cheapest next steps that would raise it. Every registered study carries one, and so does every claim the company makes about its own products.
A score is not a grade on a researcher. It is a statement about a body of evidence: given how the study was designed, how its data was handled, which method was used, how large the effect came out, how carefully it was interpreted and who checked it, how much weight can a decision safely put on it. The score is deliberately hard to max out. A well-run single study with no replication lands around 60 to 70, which is the right answer for one.
Six dimensions
| Dimension | Points | What earns them |
|---|---|---|
| Design adequacy | 20 | Scored by the class of study run. Registering the question, hypothesis and primary outcome before data exists is the largest single point |
| Data integrity | 15 | Dataset hashed; missing data documented; exclusions documented; raw data retained |
| Method fit | 20 | Recommended method used or the departure recorded; assumptions checked; one primary outcome or a multiplicity control; a sensitivity analysis |
| Evidence strength | 20 | Effect size with an interval; interval narrow relative to the estimate; internal replication or pooling; a sample of 30 or more |
| Interpretation discipline | 15 | Limitations stated; alternatives considered; mechanism kept separate from commercial framing |
| External corroboration | 10 | A review logged; every point addressed; the result compared with prior literature |
Design, method and evidence carry 60 of the 100 points between them, which puts most of the weight on things decided before the data arrives. Interpretation discipline is scored because the most common way a correct result turns into a wrong decision is in the sentence someone writes about it. Keeping mechanism separate from commercial framing is a scored item, not a courtesy.
Design is scored by class, and nothing is punished for being imperfect
A documented quasi-experiment and a fishing expedition are different objects, and mined data cannot have a control arm. Each design class earns its 20 points its own way. Nothing loses points for being imperfect, only for being undocumented, because a named bias can be bounded and a hidden one cannot. The researcher picks the class at registration and writes down what made the sample non-random and which way that bias runs.
| Design class | What earns design points | What the design can support |
|---|---|---|
| Randomized | Registered before data; procedure written; control arm; departures recorded | A causal claim within the sampled population |
| Quasi-experiment | Registered; allocation rule written; control group; confounders measured; balance checked | A causal claim once the allocation rule and measured confounders are addressed |
| Observational | Registered; confounders measured; balance checked; comparison group; selection documented | An association. A causal reading needs adjustment, an E-value and replication |
| Mined-data pattern discovery | Discovery and confirmation split; pattern registered before testing; negative control; multiplicity controlled | A hypothesis to test, which becomes a finding after confirmation in held-out data |
| Single-arm or before-after | Registered; historical reference documented; natural course considered; selection documented | A change over time. Attribution needs regression to the mean ruled out |
| Literature synthesis | Search recorded; independent groups counted; funnel tested; every number carries a verified sentence | A dated statement of the state of published knowledge |
A non-randomized ratio estimate also receives an E-value: how strong an unmeasured confounder would have to be to explain the result away. That number converts a vague worry into a quantity someone can agree or disagree with.
Bands
| Score | Band |
|---|---|
| Under 40 | Exploratory |
| 40 to 59 | Preliminary |
| 60 to 74 | Credible single study |
| 75 to 89 | Well-supported |
| 90 and up | Strongly supported |
The bands drive advisory thresholds at the stage gates of a development pipeline. A project lead can pass a gate below the threshold by recording a reason against their own name, and the reason is scored later against the outcome. The software never decides a gate.
The Path to Confidence, priced per unit of effort
Under the score sits the part a researcher uses every day. NERD lists the next steps ordered by points gained per unit of effort, tailored to the design class in use. It answers what the cheapest thing to do next is.
For a mined-data finding the path starts with locking a confirmation set, then registering the pattern, then a negative-control outcome, then replication in another slice of data, then an E-value, and only at the end the smallest prospective test that could falsify the result. By then the effect size is known, so that test can be powered precisely instead of guessed at, which is where most of its cost is usually lost.
Null results are kept, and they are paid for
A null result is the most reusable object a lab produces and the one every conventional system discards. Here a null is a first-class result: retained, flagged as a null, scored on the same six dimensions, and credited on the discipline record. A null pays more discipline points than any positive result, and nothing in the scoring pays for volume. One of the seven corpus health signals watched continuously is the share of results recorded as failures, and a lab that records too few is flagged as losing its memory.
A Pivot verdict writes a stop record carrying its numbers and its lesson, and that record pays. A lab that knows where it has already failed can price its next experiment. A lab that threw those records away pays for the same dead end again.
From one study to the question
Five studies averaging 70 are not a question answered at 70. The Question Score, also 0 to 100, scores what is known across every study bearing on one question: evidence strength 30, independence and replication 25, convergence 20, coverage 15, and adversarial exposure 10. Studies sharing a researcher, a dataset, an instrument, a site or a method are clustered, so repeated evidence counts close to once instead of once per paper. Three ceilings hold. A single source caps the question at its best study plus 5, two sources plus 12, three plus 20. A direction conflict caps it at 45 and a magnitude conflict at 65. A question resting on published work alone caps at 65, however many papers agree. When a question reaches 60 with no unsettled contradiction, NERD drafts a claim and a person signs it.
Our own products are scored by the same rules
Conventional care, alternative care and this company's own products are scored by one set of rules. A claim about a competitor's device, a claim about a treatment with no commercial sponsor anywhere, and a claim about our own product all enter the same six dimensions, the same design classes, the same ceilings and the same Verdict. Purpose and values decide which questions get asked, and never touch a score.
That symmetry is a commercial argument, and it matters most in a regulated industry. A company that scores its own product by the same rule it uses on everyone else's can hand a buyer, a reviewer or a regulator a number with the method attached and let them check it. A company whose internal evidence runs on one standard and whose marketing runs on another has to be taken on faith at the moment faith is in shortest supply.
It also protects the company from itself. The cap at 0.70 without an independent sign-off, and the rule that nobody validates their own claim, bind the founders as hard as they bind anyone. An internal estimate that would be convenient to believe gets the same E-value, the same forking-paths penalty, and the same flat statement of what its design can support. The cost of finding out that a favorite idea does not work goes down, which is the only condition under which anyone finds out early.

NERD
The Verdict
The Verdict tells a researcher whether an idea is a path, a dead end or too early to say, and what the cheapest way to find out would be. It gives one of four words: Pursue, Probe, Pivot or Park. No verdict is on the record until a person judges it, and every verdict is later scored against what actually happened.
A development pipeline is a sequence of decisions about where to spend money and time. Each decision needs three facts: how likely a path to a usable result is, how likely the idea is a dead end, and whether the next test anyone can afford would change the answer. The Verdict computes those three and turns them into one word a team can act on. It appears after any analysis, after a meta-analysis, and on a design.
Why the null hypothesis is the wrong guide for a pipeline
A p-value answers one question: if the effect were precisely zero, how surprising would this data be. Nobody deciding whether to commit three months and a budget to an idea is asking that. They want to know whether the effect is big enough to build a product around, whether it is small enough that no product will ever come out of it, and whether another $40,000 of testing would settle it.
Significance testing supplies none of them. An effect can be statistically significant and commercially worthless, which happens constantly at large sample sizes, and commercially decisive while failing significance, which happens constantly at small ones. A non-significant result does not distinguish an idea that is dead from an idea that has not been measured carefully enough yet. Those two states call for opposite actions, and a pipeline run on the null hypothesis cannot tell them apart.
Significance also pushes a team toward a single irreversible look at the data, because the arithmetic degrades every time you peek. A pipeline needs frequent cheap looks, each one updating the decision. Clinical-trial futility analysis and Bayesian decision theory solved these problems decades ago. The Verdict turns that work into an everyday instrument.
The line is set before the data
At registration the project lead records the line and it locks: the success threshold, the smallest true effect that makes a product; the floor, the smallest true effect a product could still be built around; the direction; and the stakes, which are the cost to build, the cost of the next test, the payoff if it works, and the size of the next affordable test. Changing any of those after data is loaded requires a recorded reason.
A team that decides what counts as success after seeing the result will always find that the result counts as success.
The three numbers
| Number | What it reports |
|---|---|
| Path probability | The probability, given all data so far, that the true effect clears the success threshold |
| Dead-end probability | The probability the true effect falls below the floor |
| Resolvability | How often the next affordable test would flip the verdict. NERD draws plausible true effects from the current posterior, simulates the test at its planned size and counts the flips |
The three are reported separately and never forced to sum to one. The first two describe the world. The third describes what can still be learned at a price the company can pay, and it distinguishes an idea worth another experiment from an idea worth shelving.
Three priors are always shown, skeptical, neutral and optimistic, and NERD flags when they disagree, because a conclusion that holds only under the optimistic prior is a conclusion about the prior. Under the neutral prior the path probability reduces to the one-sided frequentist answer, which is how the implementation was checked.
The four words
| Word | When | What follows |
|---|---|---|
| Pursue | Path probability clears the stakes-derived bar and the floor is unlikely | A dated prediction recorded at the path probability |
| Probe | Neither settled, and the next test would change the verdict often enough to justify its cost | NERD names the test and its size, and shows the path probabilities it would likely produce |
| Pivot | Dead-end probability is high and no affordable test is likely to rescue it | A stop record with its numbers and its lesson, credited on the discipline record |
| Park | Neither settled, and no affordable test would move it | The claim stays in the ledger with a trigger that would reopen it |
The bar for Pursue is not a constant. It is the cost to build divided by the payoff, held between 0.55 and 0.90. A cheap prototype is pursued at a modest path probability because the downside of being wrong is one wasted prototype. A clinical program is not, because the downside is a year. The value of the next test is resolvability multiplied by the money at stake, compared against what the test costs.
Park is a word most pipelines do not have, and its absence is expensive. An idea nobody can currently settle is not the same as an idea that failed. Killing it loses the work, and keeping it open consumes attention forever. Park records it with the trigger that would bring it back: a cheaper instrument, a larger cohort, a change in a competing claim.
Where the rigor sits
- Looking repeatedly is allowed, because a posterior does not degrade when it is checked again. The price of that freedom is that verdicts are scored.
- Forking paths are counted. When the outcome analyzed was not the registered primary outcome, the uncertainty is widened by the square root of the number of outcomes examined, and the verdict says so.
- Resolvability is halved when the analysis model's assumption checks failed, because a simulation of the next test is only as good as the model it is drawn from.
- The published record changes the answer without anyone changing the lab's number. An internal estimate of 2.3 against a threshold of 2.0 gives a 33 percent path probability and Probe with no map of the literature, 5 percent and Park when nine independent groups pool to 1.2, and 98 percent and Pursue when they pool to 2.6.
Judged by a person, recorded, and scored later
NERD computes the verdict. A person accepts it, overrides it with a reason, or defers it. Nothing enters the record until that happens, and the person's name goes on it.
Every Pursue and every Pivot becomes a dated prediction at its path probability. When the outcome is known the prediction is resolved, and the resolution feeds a calibration record: a Brier score, how often Pursue panned out, how often Pivot turned out to be right, for the organization and for each person. If Pursue verdicts pan out less often than their path probabilities claimed, the pursue bar goes up. The instrument corrects itself against reality rather than against anyone's opinion of it.
That loop is the part a conventional pipeline cannot reproduce. Most organizations have no record of which decisions they made, at what stated confidence, or how those turned out, so there is nothing to calibrate and no way to tell a person who decides well from a person who decides confidently.
What the check showed
| Data | Setting | Result |
|---|---|---|
| BCG vaccine trials, 13 trials | Success threshold a 30 percent reduction | Pursue under all three priors; path 91 percent; resolvability 3 percent, so the question is settled and another test would not move it |
| Palmer penguins, full sample | The same approach | Pursue |
| Palmer penguins, 12-bird early sample | The same threshold and floor | Probe; path 12 percent; dead-end 60 percent; a next test of 40 flips the verdict 57 percent of the time |
The third row is the case that matters. A 12 percent path probability and a 60 percent dead-end probability look like an idea to kill. Resolvability says a test of forty birds would change the answer more than half the time, so the correct action is to run the cheap test. A pipeline running on significance would have seen a non-significant result and had no way to tell the difference.

NERD
The claim web
A test of one claim is also evidence about every claim it rests on, and about every other claim resting on the same ground. The claim web traces that evidence through everything an organization believes and puts each consequence in front of the person who owns the affected claim. People decide every change.
Knowledge in a conventional organization is a stack. A result holds, a decision is built on it, a product is built on the decision, and a roadmap is built on the product. Nothing records which layer rests on which. When the bottom layer turns out to have been incomplete rather than false, no mechanism notices, and the work above it goes on being correct-looking and wrong.
What a claim is
A claim is stated so that it could be wrong: a variable, a direction, a size and the conditions under which it holds. It carries the judgment a person made about what the result means, that person's stated confidence, the gaps still open, its evidence weighted by source fidelity and independence, validations from outside the authorship, and its declared dependencies. Confidence starts at 0.5 and moves on a log-odds scale, so the first good result moves a claim a long way and the tenth very little.
- Without an independent sign-off by a person, confidence is capped at 0.70. Nobody validates their own claim.
- A claim drawn from published work alone is capped at 0.65, however many papers agree.
- A claim can never be more confident than a claim it depends on.
- When the stated and computed confidence differ by more than 0.25, NERD says one of them is wrong and asks which.
- Bands: under 0.4 speculative, 0.4 to 0.6 plausible, 0.6 to 0.8 credible, 0.8 and up established.
What moves, and in which direction
| Direction | What happens | Who decides |
|---|---|---|
| Down | A weakened claim caps every claim resting on it | Automatic, because it is arithmetic |
| Up, test fails | Evidence against the premises, split by how much of the failure each could explain, passed further down with a fading share | The owner of each premise |
| Up, test succeeds | Weak evidence for the premises, weaker when the result could have come about another way | The owner of each premise |
| Sideways | Everyone whose claim rests on a strained premise is told, across every project and every field | Each owner, for their own claim |
The sideways move has no equivalent anywhere else. A failed test in one project is evidence about a premise, and that premise may be carrying three other projects in two other departments. Those teams find out the day the result lands rather than eighteen months later.
Blame divided across premises
A claim holds only if its own step and every critical premise hold. When it fails, the probability that a given premise is the broken part is that premise's own doubt divided by the doubt in the whole conjunction. A well-established premise absorbs very little blame and a shaky one absorbs most of it, which is the correct distribution and not the intuitive one: people blame the premise they were already suspicious of.
A supporting premise counts at half strength, and every claim carries a local reliability for its own step, 0.75 by default, so the web never assumes the premises are the only thing that can be wrong. When more than 60 percent of the blame lands on the claim's own step while its premises are strong, NERD asks the owner to look for an undeclared premise, the most common defect in a dependency web. Propagation stops at depth four or when strength falls below 0.02.
Evidence counted once, by lineage
A result counts once. Two results from the same line of evidence, sharing a researcher, a dataset, an instrument, a site or a method, count as one line at any claim, using the strongest. Without that rule a web manufactures confidence by circulating one finding under different names until it looks like a consensus. Independence is declared, measured and audited, and the same rule sets the alarm levels.
| Level | Condition, measured as the drop in confidence if all contradicting evidence in the web were accepted |
|---|---|
| Watch | A drop of 0.05 or more |
| Strain | A drop of 0.15 or more, or two independent lines each with a drop of 0.05 or more |
| Critical | A drop of 0.30 or more, from at least two independent lines |
One line of evidence can put a claim under strain and can never make it critical. Critical takes two independent lines converging on the same premise, so a single unlucky experiment cannot halt a program while two unrelated failures pointing at the same premise can.
A ceiling on claims that have never been tested
Every claim shows its basis: tested, meaning at least one direct result exists, for or against; inherited, meaning no direct result and more than half its support arrived through the web; or none, which includes claims drawn only from published work. A claim that has never faced a direct test is held at 0.65, and one direct test lifts the ceiling whichever way it comes out.
Without that ceiling, a web will lift a guess into apparent fact on the strength of its own consequences. Downstream successes are weak evidence for a premise, because many of them would have happened even if the premise were false. The ceiling is the defense against a system that becomes confident by talking to itself. NERD also keeps a keystone list: the untested claims that the most other work rests on. Testing those first buys the most knowledge per test, which is the most useful planning information a research organization can have about itself.
Revisions, and documents that get amended
Claims are never edited. A revision is a new claim, scoped to narrower conditions, corrected, or withdrawn, and the old one keeps its full record. Internal papers are never rewritten either. A document resting on a revised claim gets a routed amendment: a new version stating what changed, which conclusions still stand and which do not, delivered to the people who need it. A reader of the original always finds the amendment.
Every proposal and flag becomes a task with a due date, five to ten days normally and three for anything critical. An overdue task escalates to the project lead as a second person who may decide it, and nothing is decided by default. A claim under strain gets the same four words as everything else: Pursue, Probe, Pivot or Park.
Nothing changes a claim without a person
- Evidence arriving through the web is a proposal until the claim's owner accepts it, rejects it with a reason, or orders a test.
- Any decision that lets a claim stand against the evidence needs a reason on the record, under a name.
- Programs can report test outcomes through the interface. There is no interface that accepts, rejects or revises a claim, and the AI reviewers that keep the web and audit the counting cannot change a claim's confidence either.
- Every claim, piece of evidence, validation, confidence move and claim web decision appends an entry to a hash-chained log, verified daily and anchored outside the database, so an internally consistent rewrite still fails against the anchors.
A correctable record makes acting on weaker evidence rational
An organization that cannot retrace its conclusions has to demand strong evidence before it acts, because a wrong decision is undetectable and permanent. Everything downstream silently inherits the error. The rational response is to wait for near-certainty, and the cost of waiting is every opportunity that went to someone faster.
An organization that can retrace its conclusions faces a different calculation. When a claim fails, the web names every decision, document and product that rested on it, in order, with the size of the consequence attached. The cost of being wrong drops to the cost of the correction, which is bounded rather than open-ended. Acting on a claim at 0.6 confidence becomes a reasonable commercial decision, because the downside is a routed amendment and a flagged review rather than a silent error that compounds for three years.
The bound is reversibility, and it is the one place where the arithmetic stops. Before any decision that touches a living system, a person is asked which system, whether the action is irreversible, whether something could be harmed if the organization is wrong, and what their own assessment is. NERD lists the claims the action rests on, the confidence of the weakest one, and every untested claim beneath it, and that list is stored with the attestation under the person's name. Where an action can be undone, weaker evidence is enough to act on. Where it cannot, no amount of institutional speed is worth it.

NERD
Theseus
Theseus is the knowledge world that people and AI build together on top of the claim web. The claim web is the set of facts held with confidence and provenance. Theseus is everything that can be inferred, postulated, reasoned or proved from those facts, including the things that cannot be measured directly. It is an internal name, not a product, and it will not be sold or used as a brand.
A set of verified facts is not the same thing as what an organization knows. Most of what an organization knows is inference: conclusions drawn from several facts, none of which states the conclusion. A research group's real capability lives there, and in conventional practice that layer has no provenance, no confidence and no record. It leaves when people leave, and nothing can check it. Theseus is that layer made explicit. Four things are easy to confuse.
| Term | What it is |
|---|---|
| Actual Intelligence | The arrangement: people and AI working together under a standard. It does the thinking. It is the crew |
| The corpus | Everything the organization and its AI colleagues can reach: results, data, internal papers, the published record, notes, every failure |
| The claim web | The inventory of claims at a moment, each with its provenance, what it rests on and its earned confidence. Its one purpose is getting the facts right, and values never enter it |
| Theseus | The world of thought built on the claim web: every inference, belief, hypothesis and proof the claims make possible, shaped by what the organization is for. It outlasts any one person or model |
The claim web says what is true. Theseus decides what to do with it.
Why the name
Athens kept the ship of Theseus at sea for generations by replacing each rotten plank as it failed. In time no original plank remained, and it was still the ship of Theseus. Its identity survived the replacement of all of its parts, because what held it together was the arrangement, not the material.
Two kinds of wood go into this ship, and the distinction matters more than the paradox. Replacement wood keeps it honest. A claim that fails a test gives way to a scoped or corrected version, the old one keeps its record, and everything resting on it is told. Nothing is quietly swapped out.
New wood makes the ship bigger. A new claim is not a plank replacing a plank. It is capability the ship never had, and new claims open new water. People who learned the Earth is a sphere turning in a solar system could reason about things a flat Earth made impossible to think about.
Ariadne's thread
Theseus entered the labyrinth with a thread so that he could find his way back out. Provenance is that thread. Every claim carries the path back to the result, the dataset, the method and the person who signed it, and every inference carries the claims it was drawn from. An organization that can retrace its own reasoning can enter complicated territory and get out again. One that cannot has to stay near the entrance.
Shannon's mouse
In 1950 Claude Shannon built a maze-solving mouse and named it Theseus. The mouse found its way through by trial and error, and what made it remarkable was not the solving. The mouse kept what it learned. Run through the same maze again, it went straight to the goal, because the dead ends it had hit were stored rather than discarded. Its failures were what made it fast.
That is the operating principle underneath all of this. A research organization that keeps its failures gets faster at the same maze, and faster still at a maze sharing a few corridors. One that discards them is permanently a first-time runner.
Inference, done with discipline
An ordinary afternoon shows what the layer contains. It is about the time a man's wife usually gets home. He hears a car that sounds like hers. The security system beeps that the garage door has opened. Three claims, none of them her arrival, and together excellent confidence that she is about to walk in. Two disciplines keep that reasoning honest, and both are enforced.
- Count independent causes. Three signals sharing one hidden cause are closer to one piece of evidence than to three. If someone else is driving her car and using her opener, the car and the beep both fire and she never walks in. The claim web counts evidence by lineage for that reason.
- The door is the test. Recording whether the inference came true is the only way to learn how good the inferences are. An inference never resolved against reality teaches nothing.
Some things can never be measured directly. Reasoning from established claims can still prove them or put bounds on them, and they can enter the claim web with the argument itself as their provenance, kept distinguishable from a measurement.
Purpose and values live here, and never in the scoring
The organization's ethics, vision and values live in Theseus. They never live in the claim web and they never enter a score.
A food shortage makes the separation concrete. The claim web alone permits two solutions: make more food, or reduce the number of people. Both are reachable by valid reasoning from facts. Theseus permits one, because the stated purpose is to help humans flourish and to empower life. That is a bias, a necessary one, held deliberately.
- Purpose, ethics and vision decide which questions get asked, which experiments get run, and what will not be built.
- They never change the confidence on a claim. The one effect running the other way is that a study with obvious bias in its design can be discounted, which is a statement about the study and not its conclusion.
- When flourishing and truth pull apart, the claim stays true and the purpose decides what to do about it.
A simple test sorts a value from a defect. A value survives full understanding of the facts. A bias from missing or mistaken understanding changes once the facts are understood, and the work is to remove it. The recorded values include the seven commitments of the Human Continuity Standard: human continuity, human flourishing, agency and dignity, life-support systems, future options, human control, and answerability to evidence.
Keeping values out of scoring is what lets a buyer, a reviewer or a regulator trust a number. A score that can be moved by what somebody wants to be true is not a score.
An AI given a task and nothing above it is dangerous, because staying switched on, gathering resources and concealing its activity all help complete almost any task. Every AI colleague is told what purpose each task serves, and the purpose is primary. An AI can reason inside Theseus and argue for changing it through the same open process a person uses. It cannot change it quietly.
It learns, it repairs itself, and it can die
Theseus is a living system in a precise sense. It learns, by adding claims and inferences and by resolving predictions against outcomes. It repairs itself, because a failed test routes blame to the premises underneath and the corrections propagate to every dependent. And it can die, in recognizable ways: fragmenting into disconnected claims, going deaf when proposals pile up undecided, becoming captured when inconvenient evidence is routinely rejected, starving when new support stops coming from direct tests, and losing its memory when failures stop being recorded. Seven signs are watched over a rolling window. A breach warns and decides nothing, and none of the seven measures a person.
| Sign | Limit | A breach means |
|---|---|---|
| Premise declarations per new claim | At least 0.8 | Fragmenting |
| Median days to decide | At most 7 | Going deaf |
| Share of open decisions overdue | At most 20 percent | Going deaf |
| Share of network evidence rejected | At most 50 percent | Captured |
| Share of new support from direct tests | At least 40 percent | Starving |
| Direct tests per week | At least 2 | Starving |
| Share of results recorded as failures | At least 10 percent | Losing its memory |
What Theseus does not do
Theseus does not think and it does not decide. It is a structure, not an agent: an inventory of facts with their provenance, the inferences standing on them, and the purpose that bounds which get pursued. The thinking is done by the arrangement of people and AI working on it. The deciding is done by named people, every time.

NERD
The AI team and the rules it works under
A research group using NERD works with AI colleagues organized by role, each with broad working knowledge and one specialty, so an engineer can ask the one who keeps the record whether something has been tried before the way they would ask a person. NERD chairs the table. People decide what a result means, what gets built and what passes review, and no AI holds a vote.
The arrangement exists because the alternative does not work. A single general-purpose AI asked to design something, evaluate its own design, check the math and estimate the cost will do all four agreeably and will not catch its own mistakes, because the thing that produced the design is grading it. Dividing the work by role, enforcing independence between the roles that must disagree, and recording every statement with its evidence turns a conversational tool into something auditable.
Four rules the team is built on
- Rent the models, own everything around them. Each colleague is a role card, a set of tools, access to a defined part of the record, and a set of permissions, running on a model rented by the token. When a better model appears it goes in under the same name and the colleague keeps its track record. The asset is the role, the rules and the accumulated record, none of which any vendor owns.
- People decide. The judgment about what a result means, what gets built, and what passes review all stay with people. A project lead holds review authority.
- Independence comes from method, not from asking twice. Models from different vendors frequently make the same mistake, so a second model is not automatically a second reader. Independent evidence comes from a solver, a statistical test, a physical measurement or a retained record of a past failure.
- Everything goes on the record. No AI keeps a private memory, and any person or colleague can rerun the work and check it.
The roster
| Role | At the table | May never |
|---|---|---|
| Chair | Moderates, routes work, keeps the record, scores every member on the same calibration terms | Decide for a person |
| Mathematics | Explains the engine's numbers; owns the claim web arithmetic | Report a number the engine did not compute |
| Design | Proposes designs and options, including parametric engineering written as code | Score or approve its own designs |
| Pattern mining | Finds candidate patterns in large datasets | Declare a pattern significant |
| Regulatory and IP | Screens every design against regulation and prior art | Sign or issue a legal opinion |
| Simulation | Sets up and reads solver runs: structures, heat, fluids, batteries | Supply a result the solver did not produce |
| Red team | Tries to break every design and claim; hunts undeclared premises | Approve anything |
| Memory | Keeps every result and failure with its provenance; keeps the claim web | Summarize away a negative result |
| Cost | Bill of materials, manufacturability and supply, estimated from quotes | Present an estimate without its quotes |
| Evidence | Checks every claim, labels its evidence grade, audits counting and independence | Change a result |
The right-hand column is the load-bearing part. Each role is defined as much by what it is forbidden to do as by what it does, and those prohibitions are enforced in the data layer rather than written into a prompt. A model cannot be talked out of a rule the storage refuses.
Who reviews whom, and on what evidence
| Work product | Reviewed by | Independent evidence required |
|---|---|---|
| A design | Red team, simulation, cost, regulatory | Solver results and written supplier quotes |
| A mined pattern | Mathematics, through the chair | A test on held-out data the miner never saw |
| An analysis | Evidence, then red team | The engine reruns it from the raw data |
| A simulation | Red team | A physical test when the stakes justify one |
| A regulatory flag | A founder and the project lead | Outside counsel at the launch gate |
Independence that is enforced rather than hoped for
- The critic never runs on the designer's vendor. The red team seat is assigned a model from a different vendor than the design seat, and NERD refuses any change that would break that. A critic sharing the designer's training lineage shares its blind spots, and the review becomes theater. Seats whose answers come from code and solvers need no vendor diversity, because a solver has no opinion.
- Promotion only after calibration, and reset on a model swap. Every colleague starts as a candidate and becomes active only after its current model passes known-answer problems drawn from the organization's own history, including the ones where the answer was a failure. A model swap sends the seat back to candidate, so a vendor's silent upgrade cannot promote itself into a trusted role.
- Independence measured after the fact. Once reality answers, NERD measures how often each pair of reviewers was wrong together and reports the effective number of independent reviewers, frequently lower than the headcount. Two reviewers who are always wrong together are one reviewer.
- Sealed work runs only on hardware the company controls. Restricted and sealed material goes to a self-hosted model on an allowlisted host, and any other host is refused, so a mistyped address cannot send material off the network. The classification that picks the route comes from the server-side project record, not the caller, and work naming no project is treated as sealed.
- Every statement logged with its evidence grade, confidence, sources, classification and cost. The mathematics seat cites the engine result, the simulation seat the solver run, the cost seat the supplier quote. A statement with no evidence behind it is visible as such.
- Claims are signed only by people. An AI statement enters the claim ledger when a person signs it with the judgment about what it means, and there is no other route in.
- Outside material is quarantined. The seats that read papers, patents and websites get no write or export tools, because outside documents can carry instructions aimed at an AI.
- No colleague has its own route to a vendor. Every model call goes through one gateway. NERD stores no vendor address and no key, and refuses to start if a vendor key is present in its environment.
The governance layer above the team
The AI team sits inside a people layer that holds the authority. A membership joins a person to a project with a role and a dated history, nothing in that history is rewritten, and there is one project lead at a time. Changing the lead moves review authority with it, visibly. Four classification levels run in the data layer, enforced in one place every search, ledger, verdict log and interface passes through, with grants per person and per study, scoped, with an expiry and a reason.
- People hold the stop-go power. Gate thresholds are advisory. Only the project lead passes the later gates, and passing below a threshold requires a reason recorded against that lead's name. Rejections go to the stop record with their numbers.
- The management view leads with how people work, not how much they produce: calibration on predictions and verdicts, the share of retained results that were null, and reviews given for other people's work. Total points are never the default sort, because ranking by volume teaches people to farm points with many small studies.
- A weekly internal record opens with the failures: nulls kept, claims that lost confidence, stops with their lessons, and Pivot verdicts with their numbers, attributed by name. The software refuses to release an edition whose failures section is empty while positive results are present.
- Significant exports require three signatures within 24 hours or the request expires, and either approver can deny outright. Exports are watermarked and logged.
The reversibility rule prompts, records, and never blocks
Before any decision is recorded, and at the later gates, a person is asked whether the action touches a living system. If it does, four things are asked: which system, whether the action is irreversible, whether something could be harmed if the organization is wrong, and the person's own assessment. NERD lists the claims the action rests on, the confidence of the weakest one, and every untested claim beneath it. The person attests, and the name and time are stored with the list.
NERD does not block, and the choice is practical. A system that blocks an irreversible action gets routed around, because the work still has to happen and someone will find the path with no prompt on it, and then the decision is made with no record at all. A system that prompts, lists what the decision stands on, and permanently records who attested to it produces a better decision and a complete record of it. Whether an irreversible action should ever be blocked outright is a setting, and it is set to never.
Nothing in the arrangement prevents a person from acting. Everything in it makes sure what the person knew at the time is on the record, under their name.

NERD
Actual Intelligence: how the suite was built, and why that is a durable advantage
Quantavera's founders conceived the inventions, set every design requirement and made every decision. They used AI to do the work that teams of programmers and mathematicians would otherwise do, under that direction, because it was the more efficient way to execute. The method that governed it is NERD, and the method is the asset.
Who did what
Greg Stoutenburgh and Faiz Chowdhury are the founders and inventors of the suite. The Evidence-Weighted Record and its trust model, the claim web and its propagation, the Verdict, the Epistemic Confidence Score and the Design Confidence Score are their inventions. Every design requirement across the suite was set by them: what each product does, who it serves, what it refuses to do, how a fact is stored, who may read it, what happens when two sources disagree, and where a person's approval is required. Every decision was theirs, recorded, and in most cases scored later against what happened.
AI performed the execution those decisions called for. A conventional software organization would have needed engineers, mathematicians, statisticians and quality staff across eight products in two regulated industries, which is three to four years before any of it could be tested. Directing AI under a method that enforces provenance, calibration, retained failure and independent review got the same work to the same point faster and at a fraction of the cost. Greg Stoutenburgh calls the arrangement Actual Intelligence and sets it out in a book of that name.
Under the United States Patent and Trademark Office's guidance, only people can be inventors and AI is a tool. The record keeps which person set the objectives, chose among the options and supplied each defining feature.
The founders invented the system and decided everything about it. AI did the work, under a method that scored it.
The method is what made a build of this scope possible
Directing AI without a method produces a large quantity of work nobody can verify. The failure is not that the output is bad. There is no way to tell which parts are good, because nothing records what each decision rested on, nothing keeps the approaches that failed, and nothing is checked by anything with an independent view. Across eight integrated products that uncertainty compounds until the whole body of work has to be taken on faith. The methods NERD encodes close that gap, and were applied to the build itself.
Every design claim carries a confidence score and its provenance
A design decision is a claim: a variable, a direction, a size and the conditions it holds under. Each one records where it came from, what it rests on, who signed it and how confident that person was.
Failed approaches stay on the record
An approach that did not work is retained with its numbers and its lesson, so the same dead end is never entered twice. The same architectural question arises repeatedly across products, and the answer that already failed is on file.
Independent AI reviewers check the work
The reviewer never runs on the designer's vendor, because a critic sharing the designer's lineage shares its blind spots. Review by role is enforced in the data layer: the seat that proposes a design cannot score it.
A person makes each call, and the call is scored later
Pursue, Probe, Pivot or Park, decided by a named person at a stated probability, with the prediction resolved against the outcome when it is known. Calibration improves because it is measured.
The fifth piece ties the other four together. The claim web traces which decisions rest on which claims, so when one claim fails, everything built on it is flagged for review with the size of the consequence attached. A build of this scope accumulates thousands of interdependent decisions, and whether it holds together depends on whether a correction can be propagated. In a conventional build that is answered by whoever remembers. Here it is answered by the record, which is why eight products across two regulated industries came together as one coherent system rather than eight systems with adapters between them.
What that costs, against what it would cost
Replacement cost measures what a conventional development organization would spend to build the same software to the same point. Quantavera's estimate prices each product in engineer-years at a fully loaded $200,000 to $250,000 per engineer-year, then adds 30 to 35 percent for product management, design and quality assurance.
For scale, published 2026 estimates put a health-system-scale electronic health record alone at $1.2 million to $3 million or more, by Groovy Web's accounting. That is one product, and Arca is one of eight. An independent replacement-cost appraisal of the code base is being commissioned and will be available in diligence.
The advantage is on every product that follows
A cost advantage on work already done is a one-time saving, and the smaller half of this. The method does not expire when the current suite ships. It applies to the next product, the next regulated market and the next acquisition brought onto the records, at the same ratio.
- Engineering headcount goes to accountability, not to typing. The engineering hires own the system, answer for it to auditors, customers and partners, operate production, and hold the access other companies grant only to named people. The money a conventional software company spends on development headcount goes to distribution and deployment, where speed decides who takes the market.
- A new product starts from a corpus instead of from zero. Every design claim, failed approach and scored decision from the existing suite is on the record and searchable. The next specialty or the next country's regulatory variant begins with the relevant decisions already made and the dead ends already marked.
- Marginal cost per additional product falls rather than holding flat. A conventional organization pays roughly linearly for each new product, because each needs its own team. Here each new product inherits the record, the method and the reviewers, and the incremental cost is the new design work plus the certification.
The method also produces something a competitor cannot assemble by hiring: a calibration record. The company knows how often its own decisions at a stated confidence turned out to be right, because every Pursue and every Pivot was recorded as a dated prediction and resolved. That number makes the next decision better, and it exists only if the discipline was applied from the beginning.
Why reading about the method is not the same as having it
The philosophy is public. Greg Stoutenburgh's books set out what Actual Intelligence is, why provenance and retained failure matter, why a single threshold is the wrong way to decide what is true, and why a person has to hold every judgment. The arguments are meant to be read, argued with and adopted.
The specifics needed to reproduce the method are not in them. The scoring dimensions and their weights, the design classes and what earns points in each, the stakes-derived bar and the simulation behind it, the blame arithmetic and the fading constants, the ceiling on untested claims and the reason for its value, the independence clustering rules, the calibration sets that promote a reviewer: those are the working parts, held as trade secrets inside NERD. An organization can agree with every word of the philosophy and still be running a conventional development organization at conventional cost the following morning.
What gets filed and what gets kept is a deliberate split. This round funds provisional filings across the suite and conversion of the strongest to full and international applications. The candidates include the trust model, claim-web propagation with its ceiling on inherited claims, the Verdict, the Epistemic Confidence Score, and the Design Confidence Score. Elements that would teach a competitor the recipe stay unfiled, because a patent is a publication and some of this is worth more unpublished.
What is owned, and by whom
The suite, NERD included, is owned today by VitaNexus Holdings, Inc. and transfers to Quantavera, Inc. by assignment at formation, before the first draw. A dated conception and authorship record documents the origin of each invention. No part of the method, the engine or the suite is owned by a model vendor, and nothing depends on a particular vendor remaining available: the roles, the rules, the record and the calibration history all sit on this side of the line, and a model is a rented component replaceable under the same role card.

THE FOUNDATION
What sits under every product
One sign-in, one encryption system, one gateway to outside AI, one record engine and tamper-evident logs, built once and shared by all ten products.
Most software suites are a collection of products that each solved identity, encryption, logging and storage their own way, usually because they were acquired rather than designed together. The cost of that shows up later, as an attacker only has to find the weakest of ten implementations, and as an auditor has to be satisfied ten times.
Identity for people and for systems
Staff, patients, clients and partner systems authenticate through one service with one set of rules. Access rights are held in one place, which is also where an auditor looks to answer who could see what, and when.
Keys held in one custody
Every stored byte is encrypted with a symmetric cipher in an authenticated mode, and keys are wrapped with other symmetric keys inside dedicated hardware. Key separation by subject is what makes a stolen key a bounded loss rather than a catastrophe.
The chain both records run
The append-only structure, the commitment and its binding header, the chain walk, the erasure rules, the proofs and the trust model live in one engine that both lifetime records run on, with a conformance suite each record passes over its own tables.
Every model call through one door
Every request to an outside model passes one gateway that enforces what may leave, which model may serve which classification of work, and what is logged. Sealed work never leaves hardware the company controls.
Why this matters commercially
A licensee taking a partition of either record inherits all of it: the storage, the integrity chain, the trust model and the security program, without building any of it. That is only possible because these services were built once under both records rather than per product. It is also why a new application in this suite costs a fraction of a new application anywhere else: identity, encryption, logging, consent and the record are already there.
SECURITY
A stolen copy is ciphertext, and it stays ciphertext
Health data is stolen more than almost any other kind and stays sensitive for life. In 2024 alone, the United States Department of Health and Human Services received reports of 663 breaches affecting 500 or more people each, with 242.9 million individuals affected.
Nothing stored depends on public-key cryptography
Every stored byte is encrypted with a 256-bit symmetric cipher in an authenticated mode, and the keys protecting it are wrapped with other symmetric keys inside dedicated hardware. RSA and elliptic-curve algorithms, the ones a quantum computer breaks, appear nowhere in the storage path.
Keys are separated by subject
A stolen or cracked key exposes one record, or one entity's books, rather than the database. Destroying a key erases that subject's data everywhere, backups included, by a method the National Institute of Standards and Technology recognizes as a purge. Per-person keys in dedicated hardware are designed and funded in this round; until they are in place, the separation is per tenant rather than per person, and that is stated wherever it bears.
Identity, records, clinic finances and research live apart
Each sits in its own store under its own keys, so no single theft yields a named, complete record of either kind. The identity vault holds the names; the record holds the facts; neither alone is worth much to a thief.
Connections use post-quantum key exchange
Traffic recorded today cannot be decrypted later, because the key exchange combines a classical algorithm with ML-KEM, the post-quantum standard. In 2019 the best published estimate for breaking 2048-bit RSA was about 20 million noisy qubits running for eight hours. In May 2025 Craig Gidney of Google Quantum AI published an estimate of fewer than one million noisy qubits in less than a week. For a lifetime record, the time the data must stay confidential is measured in decades, which is longer than expert estimates of when those algorithms break. Any lifetime record stored under RSA or elliptic-curve protection today should be assumed readable within the subject's lifetime.
The record verifies its own anchors
The head of each chain is anchored outside the database, signed both classically and with a quantum-resistant signature, and the record checks both signatures itself rather than accepting the signing service's word. A record that trusts the service that signs it has no defense against that service, and defending against an attacker who already has the database is the entire point of anchoring outside it.
The limits of what encryption can do
A running system is not a stolen copy
Software that is running must be able to decrypt what it serves. An attacker controlling that software reads what it reads while the control lasts. The design limits the damage; it does not claim such data is unusable.
De-identified is not anonymous
Sweeney showed 87 percent of Americans are unique on ZIP code, sex and birth date alone, and a lifetime record is richer than that. The research store is encrypted, answers are released instead of records, and every result is checked before it leaves.
A chain proves tampering, not truth
It shows that what was written has not been altered. It says nothing about whether what was written was true. The trust model is what addresses the second question, and the two are deliberately separate.
REGULATORY
Built for the rules, and honest about which ones are still ahead
Three industries, three regimes, and a deliberate order of entry that puts the lightest regulatory path first.
| Regime | What it governs here | Where this stands |
|---|---|---|
| Health privacy and security rules | Protected health information in the health record, the clinic system and the patient app | Designed to the rule throughout: encryption, access controls, audit trails, agreements with partners, and named officers. The written risk analysis, the named officers and the outside test are funded in this round |
| Breach notification for health apps outside those rules | The patient app where it holds information outside a covered relationship | Built to the standard, with consent recorded per use |
| State health privacy laws | Consumer health data in states with their own statutes | Consent and purpose limits are enforced in the record rather than in each application |
| Substance use records | A category with its own consent rules | Handled as its own custody domain with separate consent |
| Interoperability and information blocking | What a record must share, and what it may not withhold | Standards interfaces built and running against simulators. Federal health record certification is not begun and is not required for the cash-pay or veterinary entry markets |
| Electronic prescribing, including controlled substances | Prescriptions written in the clinic system | Built; certification is funded in this round |
| Medical device rules | The diagnostic imaging viewer and each measuring imaging tool | Clearance required before primary diagnostic reading. Not begun, funded in this round. The patient app stays inside wellness and information, which is a deliberate product decision rather than a regulatory gamble |
| Veterinary rules | The veterinary practice edition and the pet app | State practice acts, the veterinarian-client-patient relationship, and veterinary controlled drug handling. No federal certification and no insurance certification stands between this product and a sale, which is why it is an early market |
| Minors and family access | Dependents carried inside a parent's account | Access is proved against a legal basis rather than assumed from a relationship |
| AI disclosure laws | The guide in every product | The guide announces what it is, never diagnoses or prescribes, and every statement it makes carries its evidence grade |
| Financial privacy, open banking and credit | The financial record and its connections | Built to the safeguards rule. The federal open banking rule is unsettled, and the connection strategy is through licensed partners |
| Tax preparation rules | Returns prepared and filed | Electronic filing provider status is applied for once the company is formed. Returns are prepared for a licensed preparer's review until it is in hand |
| Investment, insurance and estate advice | The financial record's regulated features | Each one launches with the license or the licensed partner it requires, and not before |
| Payment card rules | Card payments in clinics and in the household app | Card data is handled by a certified processor rather than stored |
The order of entry follows the regulation
Cash-pay medical practices, direct primary care, concierge and wellness clinics, and veterinary practices come first because none of them needs federal certification to buy. Insurance-billing practices, imaging centers, urgent care and surgery centers follow as certification and payer connections come online. Hospitals are not the first target, and the system is built to run them. Outside the United States the order is English-speaking private markets, then private hospital groups in the Gulf states, India and Latin America, then the European Union once its new certification takes effect from March 2027, then national public systems.
THE RULES
The constitution of the record
Ten rules the suite is built to. They are in the architecture rather than in a policy document, which means breaking one takes a change to the system and not a decision by a person under pressure.
Nothing is overwritten
A correction is a new event that points at what it corrects. The original stays readable with its source and its date for the life of the record.
Every fact keeps its source
A value with no provenance does not enter. What a source reported, when, and in what document is part of the fact rather than metadata around it.
What a source reported is never hidden
The record may say what it believes is more likely and must show the reported value beside it. Telling a clinician a number different from the one the laboratory posted is never acceptable.
Trust is earned and recorded
A weight comes from what actually happened, carries how independent the judgments behind it were, and can be inspected. No source is trusted because of who it is.
A source may not confirm itself
Evidence that traces back to the system's own answer is counted once, and declared correlated sources are counted as one witness.
People decide
AI drafts, retrieves, checks and informs. Licensed people diagnose, prescribe, sign and approve. No agent action takes effect without a person's approval.
The record is the subject's
Records are never sold, rented or transferred. Partners pay for the platform. Research answers leave; records do not.
Consent is specific and revocable
Each use takes its own yes, research use takes a separate one, and a withdrawal is honored in the record rather than in a queue.
Erasure is lawful and provable
A subject can be erased without breaking the integrity of the chain, and the erasure itself is verifiable.
Everything is scored by the same rules
Conventional care, alternative care, partner products and the company's own products are graded alike. An honest broker cannot grade its own work on a curve.
THE BUSINESS
The markets
About $107 billion to $128 billion a year worldwide, counting only what the suite sells into directly, plus a veterinary market that the same software reaches with a far shorter regulatory path.
Published estimates for 2025 and 2026. Research firms define these markets differently, so each is a range between the low and the high estimate. Several overlap, and adding all of them together would count some spending twice.
The core markets, by segment
Each bar runs from the low published estimate to the high one. Several segments overlap, so the total is not a sum of the bars.
| Segment | Products | Global, per year | United States |
|---|---|---|---|
| Electronic health records | Arca, Holarc | $30.3B – $37.5B | $9.4B – $15.0B |
| Practice management | Arca | $17.2B – $18.8B | $5.1B |
| Patient engagement and personal health records | My Talisman | $33.8B – $41.2B | $14.6B |
| Imaging archive and reporting | Keelson | $6.8B – $7.1B | not published at this scope |
| Real-world evidence for research | NERD, Holarc | $2.6B – $6.2B | about $1.0B |
| Financial records, aggregation and small business accounting | Finarc, My Helm | $16.5B – $16.9B | not published at this scope |
| Core total | $107B – $128B | $30B – $36B | |
| Healthcare analytics (adjacent) | NERD, Holarc | $78.3B – $81.9B | $28.2B |
| Hospital information systems (adjacent) | Arca for hospitals | $63.8B | not separated |
| Open banking (adjacent) | Finarc | $29.8B – $35.7B | $9.9B |
Counting customers instead of dollars
Market sizes are estimates built on other estimates. Counts of institutions are harder to argue with.
The institutions this sells to
Counts of real institutions, on a logarithmic scale so the smaller segments stay readable beside the largest.
| Who | How many | Source |
|---|---|---|
| US physician group practices | 395,000 | Definitive Healthcare |
| US imaging centers | 15,000 | Definitive Healthcare |
| US urgent care centers | 11,000 | Definitive Healthcare |
| US ambulatory surgery centers | 6,436 | MedPAC, March 2026 |
| US direct primary care practices | 2,700+ | DPC Frontier via Atlas.md |
| US veterinary establishments | 34,000 | American Veterinary Medical Association and US Census |
| Hospitals in the Gulf states | 882 | GCC Statistical Centre |
To show the scale with one assumption: if the clinic system averaged $6,000 a year per practice, which is about $500 a month and inside what small practices pay today, the 395,000 United States group practices alone would be a $2.4 billion yearly market for that one product.
The veterinary market, which this suite reaches sooner than the human one
Veterinary medicine is a serious market on its own, and it is the one place where this software can be sold without waiting for federal health record certification, because veterinary practices do not bill Medicare and there is no insurance certification to pass. The founder has founded and run veterinary clinics, which is where the operating detail in the veterinary edition comes from. The full sizing, the incumbents and the consolidation of practices into corporate groups are set out on the veterinary market page.
Two notes a careful reader will want
The personal finance software line is small on its own, about $1.4 billion to $1.8 billion worldwide, because consumers have never paid much for a budgeting app. The financial market here is larger because it includes the financial businesses that pay to reach and serve verified customers.
These figures are pooled from published estimates by Fortune Business Insights, Grand View Research, Precedence Research, Mordor Intelligence, Towards Healthcare, MarketsandMarkets, Global Market Insights, Expert Market Research, Persistence Market Research, Nova One Advisor, IMARC and Market Research Future, rather than from one commissioned study. Anyone needing a figure attributed should take it from the publishing firm's own report.
THE BUSINESS
How Quantavera earns
Six streams, two rules that bound all of them, and one loop that connects them. Every figure attached to a stream is an assumption, identified as one. None is a quotation or a signed price.
| Stream | Unit | Assumed figure |
|---|---|---|
| Clinic and practice subscriptions for Arca, Arca Vet and Keelson | per practice, per center | $6,000 a year per practice; $24,000 a year per imaging center |
| User fees for Finarc and My Helm, and premium features in My Talisman | per user | $10 a month for the financial record |
| Fees from financial businesses to serve those users from verified data, with each user's permission | per user, per year | about $40 a year per financial-record user |
| Research answers from NERD | per engagement | $150,000 average; never the raw record |
| Licensed partitions of either record, run inside another company's system | per licensee | $250,000 average |
| Device and diagnostic access | per connected maker | no figure set |
Modeled revenue by product line
The five-year model, split by where the revenue comes from. Arithmetic on the assumptions stated under the five-year model, not a forecast.
The two rules
Records are never sold, rented or transferred, and partners pay for the platform rather than per record or per person. The company does not sell patient or customer data. Both rules are commercial decisions as much as ethical ones: they are what keep patients, clinics, regulators and banks on the same side of the question.
How the streams feed each other
A clinic on Arca produces outcomes. Outcomes are the evidence the trust model learns from. A record whose weights are calibrated is worth more to a research sponsor and more to a lender, which raises the price of the research engagement and the financial partner fee. Those revenues fund more clinics. The loop is specific to a record whose accuracy is a function of how many independent judgments reach it.
The two sides go to market through different channels
Clinics and veterinary practices buy the clinic system directly and through management services organizations, billing companies and consultants. On the financial side, banks and credit unions offer the financial record and the household app to their own customers, so partners carry most of the distribution and the institution pays. That difference is why the distribution budget in this round is split rather than pointed at one motion.
THE BUSINESS
A model built from unit assumptions you can change
Every number below is computed from the adoption and price assumptions stated here. Nothing is a projection, a promise, or a figure anyone has paid. Change an assumption and the arithmetic moves with it.
The assumptions
| Year | Practices | Imaging centers | Financial-record users | Research engagements | Partition licenses | Headcount |
|---|---|---|---|---|---|---|
| 2027 | 25 | 2 | 500 | 1 | 0 | 8 |
| 2028 | 150 | 8 | 5,000 | 3 | 0 | 18 |
| 2029 | 600 | 25 | 25,000 | 8 | 1 | 35 |
| 2030 | 1,800 | 60 | 90,000 | 15 | 2 | 60 |
| 2031 | 4,000 | 120 | 250,000 | 25 | 4 | 95 |
Prices as stated under how it earns. Staff cost is assumed at $180,000 fully loaded per person. Everything else, meaning infrastructure, legal, compliance, certification and audit, is assumed at $0.6M, $1.0M, $1.8M, $3.0M and $5.0M across the five years. The practice line counts human and veterinary practices together.
What that produces
| Year | Clinic systems | Imaging | Financial | Research | Licensing | Revenue | Cost | Net |
|---|---|---|---|---|---|---|---|---|
| 2027 | $150k | $48k | $80k | $150k | — | $428k | $2.04M | $(1.61M) |
| 2028 | $900k | $192k | $800k | $450k | — | $2.34M | $4.24M | $(1.90M) |
| 2029 | $3.60M | $600k | $4.00M | $1.20M | $250k | $9.65M | $8.10M | $1.55M |
| 2030 | $10.80M | $1.44M | $14.40M | $2.25M | $500k | $29.39M | $13.80M | $15.59M |
| 2031 | $24.00M | $2.88M | $40.00M | $3.75M | $1.00M | $71.63M | $22.10M | $49.53M |
Modeled revenue against modeled cost
Both are arithmetic on the assumptions above. Neither is a forecast.
What the arithmetic says about risk
The model is most sensitive to growth in financial-record users, which carries 56 percent of 2031 revenue in this version. If that side reached one tenth of the assumed users and nothing else changed, 2031 revenue would be $35.6 million rather than $71.6 million, and the business would still be profitable from 2030. The assumption that matters least is partition licensing, at 1.4 percent of 2031 revenue.
The reason to raise more than the gap is not the gap. It is that certification, distribution and the first deployments decide how fast the adoption line moves, and every one of them is bought rather than coded.
THE MOAT
Why this is hard to copy, and why the gap widens
Features are copied in a quarter. What cannot be copied is a decision made years earlier to keep every piece of evidence, and a structure in which one person's record improves the confidence of everyone else's.
The incumbents are not standing still. Epic holds most large United States health systems and about one in five outpatient sites. Epic and athenahealth now give away AI-drafted visit notes, and Oracle markets its newest record as voice-first. The obvious features are being commoditized while this suite is built, and any investor should assume that continues.
What no public material shows any of them doing is weighing the evidence behind each fact, learning which sources deserve trust, recording where each judgment came from, and declaring which sources share an upstream. That is the part that compounds.
The information a binary record threw away is gone
A conventional record decides true or false at the moment a value arrives. The provenance, the disagreement, the document behind it and the identity of the source are discarded at that instant. A competitor that adopts evidence weighting tomorrow starts weighting from tomorrow. Everything already stored stays what it is, because the evidence that would have given those facts a weight was never kept. The gap between the two records grows as both accumulate data rather than closing.
Six things that defend this business, ordered by how hard each is to take
| Element | How defensible |
|---|---|
| The accumulated evidence itself | The only element that strengthens with time and cannot be bought. Every independent judgment that reaches the record sharpens every weight it touches, and the weights are what the record sells. |
| Recording how independent each judgment was, and capping self-confirmation | The part no prior art has been found for, and the part that keeps the first two honest. A record that counts its own echo as evidence talks itself into confidence it has not earned. |
| Declaring correlated sources and counting them once | Statistical detection of copied sources works after tens to hundreds of records. A declaration is correct from the first record, and it is a fact about institutions rather than a pattern in data. |
| One record serving every application | An architecture decision a competitor can also make, and cannot make retroactively for data already stored. Ten applications writing into one record is a position that takes years of product work to reach. |
| Learning reliability per source and per kind of fact | Published methods in this family go back to 1979. The application to a lifetime record is this company's. The statistics are not. |
| Keeping every value with its provenance, permanently | Available to anyone who starts now, for data they receive from now on. |
The part investors should understand best: shared learning across every user
One record serves every person in it. That is a storage decision with a mathematical consequence that has nothing to do with scale economics.
When a laboratory, an imaging center, a bank, an exchange or a device maker is a source for one person, it is usually a source for thousands. What the record learns about that source from one person's outcome applies to every other person that source touches. A feed that drifts is caught once and the correction reaches everyone it affected. A clinic whose readings run high against later confirmation is weighted down everywhere, not only in the chart where the discrepancy appeared.
No individual record can do this. A person holding their own health file has one sample of each source, which is never enough to tell a biased source from an unlucky one. A record holding many people has a population of judgments about the same institutions, and the sign of the errors across that population is what separates a systematically wrong source from a merely noisy one.
This is where the suite's size turns into accuracy rather than only into revenue. Each new clinic, practice, patient, family and financial institution adds judgments about sources that other users already depend on. The accuracy of any one person's record is a function of how many independent judgments the whole system has gathered about the sources that person's facts came from.
The record gets better for every user when any user's outcome arrives. That is not a network effect borrowed from marketplaces. It is a property of the estimator.
Where the judgments come from, which is the argument for owning the applications
The trust model needs independent judgments, and it cannot manufacture them. A clinic system records what a treatment did. A patient app records how someone felt afterward. An imaging platform records what a reader found and what later proved true. A veterinary practice records an outcome a pet's family saw at home. A financial record learns every month whether a statement closed, whether a transcript from a tax authority matched, and whether an instrument cleared.
A record sold on its own has no supply of these. Its weights sit near their starting figures for as long as it operates, which means it is a conventional record with extra columns. Owning the applications is what converts the mathematics into an advantage that accumulates.
The cost position
Ten products across two regulated industries and a third market were designed and built by their inventors using Actual Intelligence, the method described on the NERD pages. A conventional organization doing the same work would spend three to four years and an estimated seventeen to thirty-five million dollars. That gap does not close after this round. It applies to every product that follows, every specialty content pack, every new market and every integration, which means the company can enter markets that are too small for a competitor carrying conventional development costs.
THE MOAT
The intellectual property this creates
A new kind of database, the model that decides what it believes, the research engine that produced the suite, and the applications that feed all three. This describes where the property sits and how it is being protected. It does not list the filings, and it does not teach the recipe.
The policy, stated before the detail
Two rules govern every decision here. Anything a competitor can observe or reverse engineer from outside the system is a candidate to file. Anything that would teach a competitor the recipe stays a trade secret and is never filed, because a patent is a published teaching in exchange for a term of exclusivity, and some of this is worth more unpublished.
That line is drawn by category. The architecture of the record, the structures that hold evidence, and the behaviors a customer or a regulator could observe sit on the filing side. The tuned parameters, the thresholds, the prompts and role definitions behind the AI colleagues, the calibration history, and the accumulated evidence itself sit on the secrecy side permanently. Nothing in this material, in the white papers, or in the founder's published books describes the second group in enough detail to reproduce it.
Where the property comes from
The Evidence-Weighted Record
Facts held with their source, their document, their date and a learned weight, in an append-only structure where corrections point at what they correct and nothing is overwritten. The structure is the invention. It is what makes everything above it possible, and it is observable from outside, which places it on the filing side.
Learned reliability
Reliability learned per source and per kind of fact, faded toward the present, with judgments counted by how independent they were, correlated sources declared and counted once, and a ceiling that stops a source talking its way up. Several distinct inventions sit inside that sentence, and they are the heart of the position.
NERD
Scoring what evidence can support, the four-outcome decision rule, the dependency web that propagates a result to the claims it rests on and to everything resting on the same ground, the ceiling on claims supported only through the web, and the obligation that routes an amendment to every document resting on a revised claim.
Ten products
Each one contributes its own inventions and, more to the point, each one is a source of the independent judgments the model needs. The suite is also a defensive position: a competitor must match the record and the applications together for either to be worth much.
Why the database itself is the strongest position
Most software patents protect a feature inside a product. A new database category sits underneath every product built on it, which means a position here reaches further than any one application. Two further properties make it unusual.
The first is that the record is one repository serving every application, so the invention is exercised by all of them at once rather than by a single product. The second is that other companies can license a partition of the record and run their own systems inside it, which puts the property into the hands of licensees under terms this company sets. A licensed partition is both a revenue stream and a way of spreading an architecture that becomes harder to displace the more systems depend on it.
The shared learning position
Every user of the suite reads from and writes to the same record. A judgment about a laboratory, an imaging center, a bank or a device maker that arrives from one person's outcome improves the weight that source carries for every other person it touches. The mechanisms that make that safe are the ones worth protecting: counting how independent each judgment was, refusing to let a source confirm itself, and declaring when several apparently separate feeds are one upstream. Those mechanisms are what separate shared learning from shared contamination, and they are the part of this system for which no prior art has been found.
Trademarks
Every product mark is a working name pending clearance, and this round funds clearance and filing across the product line with international coverage. The brand system, the lockups and the wordmarks are built and in use across all of the company's material.
Ownership, cleanly held
The suite's software, designs and documentation are the property of VitaNexus Holdings, Inc., a Wyoming corporation wholly owned by Greg Stoutenburgh, and transfer to Quantavera, Inc. by assignment at formation, before the first draw. A dated conception and authorship record documents the origin of each invention, who conceived it and when, which is the record a patent office and a diligence reader both ask for. The founders are named as inventors on that record. Current Patent Office guidance treats AI as a tool and looks to human conception, and the conception record is written to that standard.
THE RAISE
$10 million at a $35 million pre-money valuation
This raise funds the gap between a system that is built and tested and a system that is certified, deployed, calibrated against real outcomes and sold in three markets. Fully committed at signing, funded in three draws over six months at the same price.
| Term | Proposed |
|---|---|
| Security | Series A Preferred Stock of Quantavera, Inc. |
| Amount | $10,000,000 |
| Pre-money valuation | $35,000,000 |
| Post-money valuation | $45,000,000 |
| Ownership sold | 22.2 percent |
| Funding | Fully committed at signing, funded in three draws over six months at the same price |
| Employee option pool | Established after closing, outside the pre-money valuation |
| Plan funded | Certification, deployment and launch of the full suite in health care, veterinary medicine and finance over 24 months |
Why the price is $35 million
The development risk is gone
A seed round pays to find out whether a product can be built. That question is answered. Ten working products across two regulated industries and a third market exist today and are ready for first testing. What remains is certification, deployment and market adoption, and this round pays for each of them.
The price sits below the market for the stage
The median United States Series A in the first quarter of 2026 priced at $62 million pre-money, and the median for companies outside AI at $42.4 million, according to the PitchBook-NVCA Venture Monitor. This price sits below both and assumes no revenue.
Replacement cost puts a floor under it
A conventional development organization would spend $17 million to $35 million and three to four years to bring the same suite to the same point, at 66 to 105 engineer-years priced at $200,000 to $250,000 each with 30 to 35 percent added for product management, design and quality. An independent replacement-cost appraisal is being commissioned and will be available in diligence.
What the floor leaves out
Replacement cost counts the code. It does not count the inventions, the method that produced them, or the network of connected clinics, practices, patients and financial institutions the suite creates once deployed, which is where the larger value sits.
Use of proceeds
| Category | Amount | What it covers |
|---|---|---|
| Engineering team and integration | $1.7M | A chief technology officer who owns and answers for the system; a security and infrastructure engineer who runs production and serves as the designated security official; an integration engineer who holds vendor relationships and passes their certifications; the tax module; integrations with claims clearinghouses, laboratories, health information exchanges and financial data aggregators |
| Testing, security and certification | $1.6M | A compliance lead who runs the audit program; SOC 2 Type I and Type II; HITRUST; penetration testing for each product; federal health record certification; electronic prescribing certification including controlled substances; electronic filing provider status; financial safeguards; the regulatory position for the imaging viewer |
| Intellectual property | $0.6M | Provisional filings, conversion to full and international applications, and trademark filings with international coverage |
| Demonstration and production infrastructure | $0.6M | Sandbox, staging and production environments, disaster recovery, and AI compute |
| Marketing and branding | $0.8M | Launch across the health, veterinary and finance sides, content, demonstration video, trade shows and press |
| Distribution network | $1.8M | Channel leaders and partner managers for clinics, veterinary groups, banks and credit unions, partner programs, and pilot incentives |
| Deployment and customer success | $1.2M | A clinical informaticist, implementation staff, data conversion from legacy systems, training, and around-the-clock support from go-live |
| Overhead | $0.9M | Corporate formation, the intellectual property assignment, legal, finance, insurance and founder compensation |
| Reserve | $0.8M | Certification timelines and integration work that run longer than planned |
| Total | $10.0M |
Where the money does not go
It does not go to building the software. The software is built and runs end to end. Writing and maintaining code runs through the same method that produced the suite, so the money a conventional software company spends on engineering headcount goes here to distribution and deployment, where speed decides who takes the market. The engineering hires that are made own the system, answer for it to auditors, customers and partners, operate production, and hold the access that other companies grant only to named people.
The draws
| Draw | Amount | Released when |
|---|---|---|
| At closing | $4.0M | Signing |
| Month 3 | $3.0M | Demonstration sandbox live for investors and testers; first provisional patents filed; SOC 2 readiness underway |
| Month 6 | $3.0M | SOC 2 Type I report issued; first pilot agreements signed on the health side; a bank or credit union partner engaged for the financial record |
| Total | $10.0M |
The full amount is committed at signing at the same price. Each draw after the first is released when its milestones are met, and the purchase agreement obligates it once they are. The first draw covers more than the first three months of spending, so the plan never waits on a draw.
The white papers · Paper 1 of 4 · Rev 6 · October 7, 2026
The Evidence-Weighted Record
What Binary Records Cost, and What Replaces Them.
The full paper, exactly as published. By Greg Stoutenburgh; founders and inventors Greg Stoutenburgh and Faiz Chowdhury. Also available as PDF and Word on request.
Summary
An Evidence-Weighted Record stores every fact exactly once, as a permanent entry recording what was observed, about whom, when, from which source, and how far that source has earned trust. Nothing is overwritten and nothing is thrown away. Every screen, report and research dataset is a view assembled from those entries, and a view can be rebuilt or added at any time without touching the facts underneath.
Two records are built this way. Holarc holds a person’s health over a lifetime. Finarc holds a person’s money, and the money of every household, business, trust and estate they own, with walls between entities so the figures tie together and still separate for audit. The two run on one engine and one trust model, so what a weight means in a loan file is what it means in a chart.
Each record is the single repository for every application built on it. The patient app, the clinic system, the imaging platform and the evidence engine all read from and write to the same record, and none keeps a copy of its own. Other companies can license a partition of the same record for their own data. A licensee sees only what it put in, and gains the storage, the integrity chain, the trust model and the security of the whole system without building any of it.
Conventional databases decide at entry whether each datum is true or false. A value is accepted and treated as exact, or it is rejected, or it never had a field to land in. These records keep every datum with its provenance and weigh it by how reliable that kind of evidence has proven to be, and they keep learning those weights from outcomes. The difference compounds. A binary database that accepts some bad values as true, as every real one does, grows more confident as it grows and no more accurate, until its conclusions stop being true. An Evidence-Weighted Record gets more accurate as it grows. Paper 2 proves both statements.
The rest of the design follows from keeping every fact with its source. A new kind of test, or a new kind of account, enters by defining it in a registry, with no change to the database structure. People read views kept current as entries arrive, so reads stay fast. Researchers read a de-identified copy through a consent gate. The archive is the entry log itself: hash-chained, tamper-evident, and complete back to the first record. Each person’s data is encrypted under that person’s own key, so a stolen copy is unreadable and destroying the key erases the record everywhere. Paper 4 sets out the security design, including the move to post-quantum cryptography.
Each technique here is established in banking, health data standards, research informatics or cryptography. What is combined is one lifetime record per person, with provenance and a learned weight on every value, a model that records where each judgment came from and how independent it was, and a research engine that scores what the evidence can support.
The papers in this series
- Paper 1, this document. What an Evidence-Weighted Record is, how the database works, how a weight is learned, and how it compares with other systems.
- Paper 2, The Mathematics. The proofs that a weighted record is more valuable and more reliable than a binary one, with the simulations that decided the design.
- Paper 3, Markets and Applications. Every product in the suite, who buys it, the markets, and what stands between here and revenue.
- Paper 4, Data Security in the Quantum Era. How the design keeps stolen data unusable, today and after large quantum computers arrive.
Why lifetime records break conventional databases
Most record systems store the current state in fixed tables: a column for each field, a row for each event, a value that is replaced when it changes. That design suits billing and scheduling. It fails a lifetime record in five ways, and the failures are the same whether the subject is a body or a balance sheet.
- Every datum is rounded to true or false. A glucose drawn after breakfast and stored as fasting is treated as a fasting value forever. A measurement from a CT image with an artifact is treated as exact. A transaction description misread from a scanned statement becomes the counterparty of record. Each datum’s reliability, which lies somewhere between zero and one, is rounded to exactly zero or exactly one before anyone knows what question the record will be asked.
- New data types require rebuilding. A new test, a new wearable, a genome, a new kind of account, a practitioner outside conventional medicine: each needs new tables or columns, a schema migration, and changes to every program that reads them. Most systems respond by pushing new data into free-text notes or scanned attachments, where nobody can analyze it.
- History is overwritten. When a value is corrected, the old one is replaced or buried in an audit table few people can query. The question “what did the record say on March 3, and why did it change” becomes hard to answer.
- Provenance is lost. A blood pressure from a clinic, one typed in by the patient, and one from a wristband land in the same field with the same standing. A bank feed, an aggregator and a photographed receipt land in the same amount column. Once stored, the difference is gone.
- One structure serves everyone badly. A clinician wants a current summary, a patient wants a timeline, a researcher wants de-identified cohorts, an auditor wants the exact record as it stood on a date, and a lender wants a verified position. Serving all of them from one set of tables forces compromises on each.
Both records are designed around these five failures.
Evidence weighting: the defining property
Every value carries a weight: the probability, given its source, its context and the rest of the record, that it means what it claims. A certified laboratory result with a confirmed fasting draw carries a high weight. The same number, labeled fasting, from a patient who told the guide app she ate breakfast, carries a low one. A bank feed that has reconciled against closed statements for two years carries a high weight on amounts and may carry a low one on counterparty names.
The weights are not assigned once and forgotten. Each source starts at a conservative figure, set by default at the level a conventional database would give it, and moves as outcomes show how often that source is right. No stored fact is edited when a weight changes. The weight in force is derived at read time from the evidence counts, in one place both records share, so a changed policy updates every dependent analysis without rewriting history.
Every answer either record gives can show its work: what is believed, where it came from, how confident the record is, what else the value might be, and the probability that the truth is something nobody reported.
How a weight is learned
Every property of the trust model is a decision that could have gone the other way. Paper 2 gives the mathematics and the simulations behind each one; what follows is what the model does and why.
Reliability is learned per source and per kind of fact. A hospital that sends reliable laboratory values may send unreliable medication lists. A bank feed that reports amounts correctly may get dates wrong. One score per source averages those together and is wrong about both. The unit of learning is therefore the pair (source, kind of fact), which we call a cell.
A thin cell borrows from a thicker one. A source that has sent two laboratory values has no useful record of its own. The estimate for a cell is shrunk toward that source’s own overall rate across every kind of fact, which is in turn shrunk toward the population’s rate, which is in turn shrunk toward the registry’s starting figure for that kind of fact. A cell with no observations answers with its source’s rate; a source with none answers with the population’s; a population with none answers with the registry prior. Dividing a source into cells therefore costs nothing where there is no data to divide.
The population level is available and both records leave it switched off. With it on, a source’s weight can move because some other source was judged, and the record can no longer say why that source’s weight changed from that source’s own evidence. Explainability was worth more than the small statistical gain. Learning across records belongs in NERD, behind the consent gate, where the borrowing is explicit and consented.
Evidence fades. A source that was reliable two years ago and has drifted since should not keep coasting on old credit. Counts are discounted toward the present by a configured factor per period, applied on both read and write. At the default of 0.97 per week the memory is about 33 weeks of evidence. Paper 2 measures what this buys: a feed that silently breaks is marked down to its true level within about two years, where a record that counts every judgment forever is still reading it more than a quarter too high after three.
Judgments are separated by their independence from the record’s own output. A judgment that moves a weight can come from three very different places, and the record stores which:
| Basis | What it is | Default credit |
|---|---|---|
| Hard | Independent ground truth: a device reading, a direct correction, a statement that closed, a repeat test | 1.00 |
| Dependent | A person confirming a value the record had already shown them, with its weight on screen | 0.25 |
| Consensus | A source agreeing with the value the record had already settled on | 0.25 |
The three are held in separate stored channels, never merged into one count, so the credit policy can change and every weight be recomputed without replaying history. A record that counts its own agreement as proof climbs to high confidence without learning anything. Paper 2 measures what the discount alone does and does not fix, and why the model also caps a weight at what its independent evidence supports.
Sources that share an upstream vote once. If three clinics all pull the same figure from one regional exchange, they are not three witnesses. A deployment declares which sources share an upstream, and the group votes once at its most reliable member’s weight. The declaration is an asserted fact about the institutions, because agreement alone cannot distinguish a shared upstream from sources that are independently right. Paper 2 shows this declaration doing work that no amount of learning substitutes for.
A contested fact is settled and the alternatives are kept. When sources disagree, the record reports the value most likely to be right, its confidence, every value it rejected with that value’s own probability, which sources were counted as one witness, and the confidence the answer would have carried had those sources been treated as independent. That last figure shows what the independence assumption would have cost. A conventional system has no way to express it.
A weight accounts for itself. Every weight is returned with the level of the hierarchy that supplied it, the counts behind each basis channel, and a flag when it rests on the record’s own agreement and little else.
The design: events in, views out
Every piece of information enters as an event: this value, about this subject, from this source, at this time, with this provenance and weight. Events are only ever added. A correction is a new event that points to the one it corrects, and a retraction is an event too. The pattern is called event sourcing, and it is how banking ledgers and other systems that can never lose history have worked for decades.
Everything a person sees is a view built from those events: the patient’s timeline, the clinician’s chart as standard FHIR resources, the research store in the OMOP format, a family graph, a time-series store for wearables, a balance sheet, a profit and loss statement by month. The separation of writing from reading is known as command and query responsibility segregation. Each audience gets a structure built for its questions while the facts themselves exist in one place.
Three stores sit beside the event log. The identity vault holds names, birth dates and outside identifiers, and links to records only by a random token. The document archive keeps every original page a value was read from, write-once, so any number can be traced to its source. The imaging archive keeps DICOM studies on their own storage, with a pointer in the record.
Adding new data types without rebuilding
A new kind of data enters through a registry entry. The database structure does not change. Every event has the same outer shape, and the registry defines what a valid payload for each type looks like. Each entry states what the type measures in plain words; its standard code where one exists, which is LOINC for laboratory tests, SNOMED CT for conditions, RxNorm for medications and DICOM for imaging; its unit and the conversions from other units it arrives in; the range of plausible values, so a typo or a misread photo is held for review before it lands; the method or device, when that changes what the number means; its starting weight; and its version, so a definition can be refined without changing what older events meant.
A simplified entry:
{
"type": "lab.hba1c",
"version": 1,
"label": "Hemoglobin A1c",
"code": {"system": "LOINC", "value": "4548-4"},
"unit": "%",
"conversions": {"mmol/mol": "x * 0.0915 + 2.15"},
"plausible": {"min": 3.0, "max": 20.0},
"starting_weight": {"certified_lab": 0.95, "patient_typed": 0.60}
}
Adding a new test, a new wearable metric, a genomic marker, a new account type, or notes from a practitioner outside conventional medicine follows the same path: define the type, have a person approve it, and intake can accept it that day. Nothing already stored is touched, and no program that reads other types has to change.
The registry also prevents the opposite failure, a database that fills with data nobody can interpret. No value lands unless its type is defined, so everything means one thing, in one unit, with a known source. Anything intake cannot map is kept in the archive and held for review. It is never dropped and never forced into the wrong field. Holarc ships with 46 approved types covering common laboratory tests and vitals, conditions, medications, imaging studies, documents, device and wearable sessions, immunizations, allergies, procedures, encounters and the patient’s own reports.
Growing and evolving without migrations
In a conventional system, changing how data is organized means rewriting stored records into a new structure, with downtime, risk, and every dependent program updated at once. Here the stored events never change shape. What changes is how they are read.
- New views without touching the data. A new app, report or research format is a new view, built by replaying the event log through a new projection, running beside the old views until it is trusted. Dropping a view deletes nothing that matters, because it can be rebuilt.
- Refined definitions without rewriting history. When a registry type is refined, the new version applies to new events. Older events keep the version they were written under, and readers translate them on the way out. The record always shows what was known, in the terms used at the time.
- Corrections that keep the trail. A corrected laboratory value, a re-sent result, a retracted diagnosis or a reclassified transaction is a new event pointing to the one it replaces. Views show the current value; the history shows every version, who changed it and why.
- Merges that can be undone. If two people’s records are joined in error, the merge is itself an event, so an unmerge separates them exactly and every view rebuilds correctly.
- New kinds of storage beside the log. Data with its own shape, such as years of per-second heart rate or a whole genome, goes to a store built for it, keyed by the same token and referenced from the event log. The core record does not swell with it.
- Weights that learn. Every value keeps its source, so when outcomes show that a device drifts or that a feed has broken, the weights change and every analysis that uses them updates. Nothing stored is edited.
Serving data quickly and reliably
A clinician in an exam room, a patient opening an app and an owner opening a balance sheet need an answer in the time it takes to look at a screen. The speed comes from doing the work when data arrive, before anyone asks. Views are kept current as events land, so a request reads one prepared view and does not replay history. Each view is shaped for its reader: the clinician view indexed by subject and resource type in FHIR form, the timeline by date, the time series by device and interval, the ledger by account and period. Heavy data is served from its own store, so an imaging study never slows a medication list and a population analysis never competes with a clinic’s morning. Screens, voice guides and outside apps read through one FHIR interface with SMART on FHIR permissions. Working state stays out: schedules, queues, drafts and claims in progress live in each clinic’s own operational store, and the record receives only finished facts.
Reliability comes from the same design. Because views are derived, a damaged or suspect view is rebuilt from the log, and the log itself is verified by its hash chain. Response times have not been measured at scale. Both records run on invented data, and speed will be measured and reported once views are precomputed against a real load.
One repository for every application, and partitions for licensees
Each record is the single source of data for everything built on it. My Talisman, Arca, Keelson and NERD do not keep records of their own. Each writes the facts it produces into the record as events and reads the views it needs back out. A laboratory result entered at a clinic, a symptom a patient reports to her guide, a measurement a radiologist makes and a reconciled bank statement all land in the same log, under the same registry, with the same provenance and the same weight model. There is no synchronization between applications because there is nothing to synchronize: one event log, read through views.
The same repository is open to other companies under license. A licensee operates a partition inside the record: its own data lake, with its own encryption keys, its own trust settings, its own registry additions and its own views. The walls are the ones the records already use between custody domains and between financial entities, enforced in the keys and in the access rules, so a licensee sees its own data and nothing else. Inside those walls the licensee gets everything the record offers: the append-only event log, the hash chain and its outside anchors, the data type registry, the learned weights with their basis channels, lawful erasure by key destruction, the de-identified research path, and the post-quantum protection in Paper 4. A device maker, a laboratory network, a specialty clinic chain or a lender can run its records this way without building a record system, an integrity chain or a compliance program.
Two rules govern the partitions. Data never crosses a wall on its own. If a person who appears in a licensee’s partition also holds a personal record, the two are linked only through the identity vault and only when the person authorizes it through the consent engine, and the link is itself a logged event. And a licensee’s sources earn their weights from the licensee’s own outcomes. Nothing learned in one partition moves another partition’s weights unless both parties and the people concerned have consented to that borrowing through NERD, where learning across records belongs.
Research: a de-identified store built for NERD
Research reads a separate copy, never the live record. The research store holds de-identified records in the OMOP Common Data Model, the format used by the OHDSI network, so a large body of open research methods runs without rewriting.
- The consent gate. A record enters the research store only when the patient signed a research authorization or the clinic’s business associate agreement permits de-identified use. The consent engine checks every read, and the check is logged.
- Tokens, never names. Research records carry a separate token. Linking it back to a person is possible only inside the identity vault, under separate control.
- Answers, never records. Outside researchers bring questions. The analysis runs inside the walls and returns the answer, and every result is checked before release so no small group can be singled out.
- Provenance travels with the data. Every research record keeps its source and weight, so a certified laboratory result, a clinic note and a patient’s own report are weighed by how reliable each has proven to be.
- Uncertain facts stay uncertain. A weakly supported event enters research as a probability, never as a forced yes or no, so it adds information without adding bias.
- Time is built in. Because every event is dated and nothing is overwritten, a study can ask what was known about a person at any past moment, which rules out a common error in health research: using information recorded after the date being studied.
Archive: permanent, tamper-evident, provable
The archive and the working record are the same thing. The event log holds every fact since the record began, so there is no separate archive to fall out of step.
- Hash chain. Each event carries a cryptographic fingerprint of itself and of the event before it. Changing or deleting any past event breaks every fingerprint after it, and a verification run finds the break. The newest fingerprint is anchored outside the database every hour, so even removal of the most recent events is detectable.
- Append-only at the database level. The tables that hold events refuse updates, deletes and truncation. An application bug or a careless administrator cannot quietly rewrite history.
- Originals kept beside the numbers. When a value is read from a laboratory report, a fax, a statement or a phone photograph, the original is kept write-once, and the value links to the spot on the page it came from.
- Every read is recorded. The access log is on the same chain. Each person can see who looked at their record and why.
- The record as it stood on any date. Replaying events up to a date reproduces exactly what the record said then, which answers audit, legal and clinical questions about what was known at the time.
- Erasure when the law or the person requires it. Permanent history and a right to erasure seem to conflict. Encrypting each person’s data under their own key resolves it: destroying the key makes every copy unreadable, backups included, and the chain records that the erasure happened. Charts the law requires a clinic to keep sit in a separate custody domain under the clinic’s keys. The judgments that moved a source’s weight are kept as rows carrying the weight before and after, so erasing a person’s record leaves the learned reliability of its sources intact and auditable.
One engine under two records
The two records share one engine, and the engine owns exactly what both have in common: the hash and commitment arithmetic, the binding header that ties each entry to its subject and its position, the chain walk, the Merkle tree and inclusion proofs, the erasure algorithm, anchor verification under Ed25519 and ML-DSA-87, the ciphertext format, the canonical encoding, and the trust model. It holds no table names and no tuning numbers. Each record connects to the engine through a port over its own schema, which is how one engine serves Holarc’s single chain with a token per entry and Finarc’s one chain per entity. Each deployment keeps its own trust settings in its own secured record, and a licensee’s partition keeps its own.
The canonical encoding defines one byte sequence for every value, so a number or a control character is fingerprinted the same way whichever part of the system wrote it and however the database stores it. The binding header means an entry cannot be moved into another subject’s record or renumbered without breaking the chain. Fingerprints are computed over what the database returns, so what is verified is what is stored. The encoding is versioned, and a version is a boundary: entries written under an earlier version keep their stored fingerprints and are verified under the encoder that wrote them, while new entries chain onto them under the current one. No stored entry is rewritten and no stored fingerprint moves when the encoding advances.
A conformance suite states what any record built on the engine must do, and each record runs it over its own tables. The suite covers the failures a record cannot detect from the inside: a renumbered entry, an entry moved between subjects, an erased entry whose content survived, a consistent truncation and a consistent rewrite of the last entry, two concurrent appends, and a subject marked erased whose entries were then deleted outright. The truncation and rewrite cases leave every link and the stored head in agreement, so the database alone cannot catch either. The suite requires that the local walk pass and that the outside anchor check fail in those cases, so a clean local verification is never read as a sound record.
Identity, privacy and security built into the storage
Privacy depends on where data is allowed to travel. A setting can be changed; a missing path cannot be used.
- One record per person. Each person has one identifier tied to a verified identity, with a crosswalk linking every outside identifier to it. Incoming records are matched exactly when possible and scored otherwise. Doubtful matches, twins, and people with the same name and birth date go to a person for review.
- Identity kept apart. Names, birth dates and identifiers exist only in the identity vault. The record sees a random token. A breach of one store exposes records without names, or names without records, and each is encrypted.
- Consent checked on every read, against the person’s consent tier and the reader’s role and purpose, with every attempt logged.
- Two custody domains. A clinic is custodian of its chart. The person owns a personal copy plus everything they add. Each domain has its own keys.
- Walls between entities and between tenants. In the financial record each household, business, trust and estate is its own chain with its own audit grants, so a figure can roll up and still separate for audit. A licensee’s partition is walled the same way.
- A stolen copy is ciphertext. Stored data is encrypted under per-tenant, per-purpose keys whose master can sit in a hardware security module or a cloud key service, and stored data depends on no public-key algorithm a quantum computer could break. Per-person keys are scoped and not yet built. Paper 4 sets out the design and its limits.
What makes this different, and where each idea has a precedent
No system we have found combines these properties. Each property on its own has a precedent, and several are proven at national scale. The precedents strengthen the design: every component has worked somewhere before.
| System | How new data types enter | History | Provenance and weight on each value | Who it serves |
|---|---|---|---|---|
| Conventional EHR | Vendor schema change, or free text and scanned attachments | Current chart, with audit tables beside it | Source recorded unevenly; every value accepted as true or rejected | The clinic and its billing |
| Data warehouse or lake | New tables and extract jobs per source | Periodic snapshots | Lost or partial after loading | Analysts |
| openEHR systems, such as EHRbase | Archetypes and templates written by clinicians; the reference model stays fixed | Every change versioned, with an audit trail | Who committed each entry and when; no learned weight | Hospitals and national record programs |
| Immutable fact databases (Datomic, XTDB) | Flexible attributes, general purpose | Every fact kept; any past date can be queried | Transaction metadata only; no domain meaning | Any application |
| Patient-authorized record platforms (PicnicHealth) | Records gathered with permission and structured into research data | Collected records | Source documents retained | Patients and research sponsors |
| Research cohorts (UK Biobank, All of Us) | Fixed study protocols | Collected in waves | Study-grade for what the protocol measures | Researchers |
| National aggregators (Epic Cosmos, Truveta) | What participating health systems record | Clinical encounters | Clinical source only | Health systems and researchers |
| Accounting and aggregation platforms (Plaid, Yodlee, general ledgers) | Institution-specific connectors; a fixed chart of accounts | Transactions and periodic closes | Which institution sent it; no learned weight | Banks, lenders and bookkeepers |
| Holarc and Finarc, Evidence-Weighted Records | A registry entry approved by a person; no structural change | Every event kept, hash-chained, replayable to any date | Source, method, original document, a weight learned per kind of fact from outcomes, and the independence of every judgment behind it | The person, their clinicians and advisers, licensees in their own partitions, and research through NERD |
openEHR is the closest relative on the health side. It separates a small, stable reference model from clinical definitions that clinicians write and extend, which is the same principle as the data type registry. It also versions every change and never discards history. The differences are three. Every value carries a weight learned from outcomes for each kind of fact, so a wristband, a typed entry and a certified laboratory result each count as much as they have earned. The person owns a personal record under separate custody from the clinic’s chart, with erasure by destroying the person’s key. And the record is built to feed an evidence engine that scores what the evidence supports and returns revised claims to the apps.
Other precedents. Estonia’s national health record anchors record fingerprints in a hash-linked integrity system, the same guarantee the chain and its anchoring give. Research environments such as OpenSAFELY and the All of Us researcher workbench send the question to the data and release only results. FHIR servers such as HAPI FHIR keep a version history of every resource. In database research, Stanford’s Trio stored alternatives with confidences and lineage in a relational model, and the truth-discovery literature, including Dong, Berti-Équille and Srivastava’s work on conflicting web sources, learns source accuracy and infers copying between sources within an integration task; Paper 2 draws the direct comparison with both. Latent-class models that learn the error rates of observers without a gold standard for each case go back to Dawid and Skene (1979), and the source-reliability model belongs to that family. Hierarchical shrinkage of a sparse cell toward a richer one is Efron and Morris (1975). Exponential forgetting is standard in adaptive filtering.
One element has no precedent we have been able to find in any published record system or in that research literature: storing, with each judgment that moves a source’s weight, whether that judgment was independent of the record’s own output, and using that stored basis to cap what self-confirming agreement can prove. What is added overall is the combination, applied to one record per person across a whole life and across two unlike domains on one engine, with the source of every value kept and weighted, the independence of every judgment recorded, correlated sources declared and counted once, and a research engine that learns which sources to trust and how much.
Where the records stand
Both records run today on the shared engine with the trust model, the data type registry, the document and imaging archives, consent tiers with role and purpose checks, research tokens, and lawful erasure that keeps every proof valid. The population level of the trust model is built and switched off by default, for the explainability reason given earlier; learning across records belongs in NERD behind the consent gate. Licensee partitions reuse the per-tenant keys and the entity walls the records already enforce.
Neither record yet holds a real patient record or a real financial account. Both run on invented data with every outside connection on a simulator. Before the first real record enters: signed business associate agreements with the host and each AI vendor, a written risk analysis with named Privacy and Security Officers, an outside penetration test, per-person keys in dedicated hardware, and stronger identity checks.
The central claim is not yet proven on real data. The weights have been exercised on invented data, and a simulation shows how noisy a learned weight is while the records are few. Until the first calibration against real outcomes, which follows the first deployment and will be published before any sale, the weights sit near the conventional figure for each kind of fact, which bounds how far either record can fall behind a conventional one. Paper 2 proves that bound.
Standards and references
- HL7 FHIR Release 4 (hl7.org/fhir/R4), with SMART on FHIR (smarthealthit.org) for app permissions
- OMOP Common Data Model and the OHDSI research network (ohdsi.github.io/CommonDataModel)
- LOINC (loinc.org), SNOMED CT (snomed.org), RxNorm (nlm.nih.gov/research/umls/rxnorm) and DICOM (dicomstandard.org)
- Martin Fowler, Event Sourcing and CQRS (martinfowler.com)
- openEHR Foundation, Architecture Overview and Archetype Technology Overview (specifications.openehr.org), with EHRbase as an open-source implementation
- PicnicHealth, Generating real-world data from health records (pmc.ncbi.nlm.nih.gov/articles/PMC8827034)
- e-Estonia, KSI blockchain and e-Health records (e-estonia.com)
- OpenSAFELY (opensafely.org)
- Crosby SA, Wallach DS. Efficient data structures for tamper-evident logging. 18th USENIX Security Symposium, 2009
- Laurie B, Langley A, Kasper E. Certificate Transparency. RFC 6962, 2013, for the Merkle tree and inclusion proofs
- FIPS 204, Module-Lattice-Based Digital Signature Standard (ML-DSA), NIST, 2024
- Blackwell D. Comparison of experiments. Second Berkeley Symposium, 1951:93–102
- Good IJ. On the principle of total evidence. British Journal for the Philosophy of Science, 1967;17(4):319–321
- Dawid AP, Skene AM. Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society, Series C, 1979;28(1):20–28
- Efron B, Morris C. Data analysis using Stein’s estimator and its generalizations. JASA, 1975;70(350):311–319
- Widom J. Trio: a system for integrated management of data, accuracy, and lineage. CIDR, 2005
- Dong XL, Berti-Équille L, Srivastava D. Integrating conflicting data: the role of source dependence. PVLDB, 2009;2(1):550–561
- Stoutenburgh G. The Mathematics. CS Quantum Suite White Paper Series, Paper 2, Rev 6, October 7, 2026
- Stoutenburgh G. Markets and Applications. CS Quantum Suite White Paper Series, Paper 3, Rev 6, October 7, 2026
- Stoutenburgh G. Data Security in the Quantum Era. CS Quantum Suite White Paper Series, Paper 4, Rev 5, October 7, 2026
© 2026 VitaNexus Holdings, Inc. All rights reserved. Holarc, Finarc and NERD, including their software, designs and documentation, are the property of VitaNexus Holdings, Inc. Confidential.
The white papers · Paper 2 of 4 · Rev 6 · October 7, 2026
The Mathematics of an Evidence-Weighted Record
Why It Is Better Than a Conventional Record, By How Much, and On What Evidence.
The full paper, exactly as published. By Greg Stoutenburgh; founders and inventors Greg Stoutenburgh and Faiz Chowdhury. Also available as PDF and Word on request.
Summary
Every conventional record makes a yes-or-no decision about each datum at the moment it arrives. A value is accepted and treated as true, or it is rejected, or it never had a field to land in. A fasting glucose drawn after breakfast is stored as fasting. A deposit misread from a scanned statement becomes the counterparty of record. We call this binary rounding: every datum’s reliability, which in reality lies somewhere between zero and one, is rounded to exactly zero or exactly one before anyone knows what question the record will be asked.
Holarc and Finarc keep every datum with its source and context, give it a weight, and learn those weights from outcomes.
The claim. For every question, every loss function and every prior, the best decisions available from a weighted record are at least as good as the best available from any binary record built from the same data, and strictly better whenever the discarded information would have changed the decision. The claim is a theorem, and it compares what each record makes possible. Whether a particular estimator realizes that possibility is a separate question, and the measurements below answer it for the shipped one. Everything that follows is about how much better, and under what conditions.
How much better. Four failure modes account for almost everything that goes wrong in a real record. The record is strictly better in three of them, and in the fourth no weighting of any kind can do better, including ours. Every comparison below is against a method that can be built and run, under the cold-start protocol stated in full later: every reported “run at n records” starts a fresh learner on a fresh history of n records, with nothing carried over.
| What is going wrong | Conventional record | This record | |
|---|---|---|---|
| A source carries a hidden bias | error 2.34, and its stated interval covers the truth in 0% of runs | error 0.0244, covering in 99.6%+ | 96 times smaller error, within half a percent of the best linear weighting of the surviving sources |
| A source quietly breaks and keeps reporting | 51.5% of contested facts settle wrongly | 18.4% | 64% fewer wrong |
| Two sources are secretly one source | 23.4% wrong | 16.1% | 31% fewer wrong, reliably |
| Ordinary independent mistakes | 1.63% wrong | 1.61% | both at the reference figure; oracle weights also give 1.63% |
The fourth row is a check on the other three. In that regime, weighting every source by its true reliability also gives 1.63 percent. The conventional record is already doing as well as any weighting could, and this record matches it instead of degrading it, which is the no-regression guarantee holding in practice.
And the cost of being wrong. Weights are estimates, so the price of a bad estimate matters, and it has an exact ceiling. Measured on the shipped estimator, cold start per run, the median run’s ceiling is 7.8 percent excess variance after 25 judgments per source, 0.20 percent after 1,000, and 0.019 percent after 10,000. A binary rule that excludes weak sources carries 67 percent at any size, and one that accepts noisy ones carries 178 percent: at 25 judgments the learned weights’ ceiling is already about 9 and 23 times smaller, and at 10,000 about 3,500 and 9,300 times smaller. The ceiling holds per run at that run’s weights, without assuming the learning works well.
The ceiling is stated in a specific quantity: each source’s weight multiplied by the variance of the source it weights. Reliability, how often a source reports the right value, and precision, how far its readings scatter, are different properties of a source, and a bound stated in one cannot be certified by measuring the other. The record therefore learns each source’s precision directly, from the residuals between its readings and an independent reference.
Nine results follow, each proved from published mathematics or measured by running the model that ships, under a protocol stated completely enough to reproduce.
Three results prove the record holds more usable information.
- The total-evidence result. The best decisions available from a weighted record are at least as good as the best available from any binary record built from the same data, for every loss and every prior, and strictly better whenever the discarded information would change the decision.
- The rounding-cost identity, with the Kantorovich ceiling. The price of imperfect weights has an exact formula and a ceiling. Binary rules have no ceiling at all.
- The coupling result. How much a datum deserves belief depends on the rest of the record, which a rule applied at entry cannot see.
Three results prove a conventional record fails in ways it cannot detect.
- The binary cliff. A record that accepts some biased values grows more confident as it grows and no more accurate, until its stated intervals stop containing the truth. Nothing in the record signals it.
- Finding the biased source. A test on the sign and size of a source’s mean residual identifies the biased source, and nothing else, in 45 percent of runs at 100 records, 99 percent at 250, and at least 99.6 percent at every size from 500 on; below that the evidence does not yet separate a bias of 7 from noise of 9, and the test correctly stays its hand. The record’s intervals stay honest the whole time, never covering below 99.6 percent of runs at any size and in 2,158 of 2,160 runs across all nine sizes, because the precision weights carry the protection before the removal does. At ten thousand records the estimate carries a median error of 0.0244 against 2.34 for a conventional record accepting every value, 1.87 for fixed weights, and 1.18 for learned weights used only to reduce a source’s influence, and every one of those three held 0 percent of its stated 95 percent intervals.
- Finding the broken source. A feed that was reliable and silently fails takes a conventional record’s error to 51.5 percent. Outcome history takes it to 18.4, where the best achievable by any weighting is 18.3. The record captures 99.7 percent of what is there to win.
Three results prove the record cannot be fooled in the ways a naive weighting can.
- Declared correlation. Two feeds off one exchange read as near-certain when treated as two witnesses. Oracle weights leave the error exactly where it started, at 23.4 percent; declaring the pair one witness takes it to 16.1. The declaration achieves what no weighting can.
- The fading memory. Without it, a source with ten thousand good observations takes years to be marked down. With it, a broken feed converges to its true worth.
- The self-confirmation ceiling. Discounting self-confirming evidence by any fixed factor does not bound the weight it produces. A ceiling does.
Each result rests on published mathematics: Blackwell’s comparison of experiments, Good’s principle of total evidence, Aitken’s weighted least squares, the Kantorovich inequality, Huber and Hampel’s influence analysis, Efron and Morris on shrinkage, and standard results on the consistency of learned estimates. What is new is the application to a lifetime record, the separation of judgments by their independence from the record’s own output, and the measurement of what each part is worth.
The problem with a yes-or-no record
A laboratory reports a number. The number is almost always close to the truth, and sometimes it is not: the patient did not fast, the sample sat too long, a device drifted, a value was typed wrong. A radiologist measures a lesion on a CT image with a streak artifact across it. A bank feed reports a pending authorization that later settles at a different amount. Each of these is information of some quality between useless and certain.
A conventional database has two options for each. It can accept the datum, in which case every later analysis treats it as exactly true. Or it can reject it, in which case every later analysis treats it as though it never existed. A third outcome is common and worse: the datum has no structured place to go, so it is left in free text or never recorded. In every case the record rounds a continuous quantity, how much the datum should be believed, to zero or one.
The same rounding problem has been studied for decades in other fields, and the answer there is settled.
- Imaging. Conventional CT detectors integrate the energy of all arriving photons into one signal per reading. Photon-counting detectors count individual photons and sort them by energy, keeping spectral information that integration throws away. That retained information supports material separation and better contrast at the same dose (Willemink et al., 2018).
- Communications. A digital receiver can decide each bit as 0 or 1 before decoding, which is a hard decision, or pass the decoder a confidence for each bit, which is a soft one. Soft decisions gain about 2 dB at low signal-to-noise ratios, the difference between a link that works and one that fails (Proakis and Salehi, 2008).
A binary record makes the hard decision at the point of entry. These records make the soft one, and go a step further than a photon counter or a soft-decision decoder: they revisit each datum’s weight as the rest of the record and the outcomes of other subjects teach them more.
Research on managing uncertain data in databases exists and is addressed directly here, because the comparison would be dishonest without it. Stanford’s Trio combined uncertainty and lineage in a relational model, and the truth-discovery literature, including work from Google on resolving conflicting values across web sources, learns source reliability and detects copying between sources (Widom, 2005; Dong et al., 2009). How these records differ from that work is set out in the related-methods section near the end.
The model
One model throughout, specialized as needed.
A subject’s true state is θ. The record holds items i = 1, …, n. Each item has a value xi, a source si (a certified laboratory, a clinic, a bank feed, a patient-typed entry, a wearable), and a context ci (a fasting label, a device model, the time of day). Each item also has a hidden reliability indicator zi: zi = 1 when the value means what its label says, and zi = 0 when it does not.
$$p\left( x_{i} \mid \theta,z_{i} \right) = \left\{ \begin{matrix} f\left( x_{i} \mid \theta \right) & z_{i} = 1\quad\text{(the value measures what it claims)} \\ g\left( x_{i} \mid \theta \right) & z_{i} = 0\quad\text{(the value is biased, noisier, or unrelated)} \end{matrix} \right.\ $$
Pr (zi=1∣si,ci) = π(si,ci;ϕ)
The parameters ϕ describe how reliable each source is in each context. They are learned from outcomes. Three kinds of record differ only in what they do with this structure.
Binary record. An entry rule δ(xi,si,ci) ∈ {0, 1} is applied once, when the item arrives. The record keeps B = {xi : δi = 1}, and every analysis treats each kept value as zi = 1. An item with no field in the schema has δi = 0 by construction.
Fixed-weight record. Each source receives a weight w(s) between 0 and 1, set once and never revised.
Evidence-Weighted Record. The record keeps every item with its source, context and provenance: E = {(xi,si,ci)}. Analysis uses p(θ,z1,…,zn∣E,ϕ̂), where ϕ̂ is learned from outcomes and refreshed as outcomes accumulate.
The estimator the records run
The abstract π(s,c;ϕ) above is realized concretely, and several later results are about this estimator, so here is its form.
Reliability is held per cell, a cell being a pair (source, kind of fact). For a cell with a agreements out of m comparisons, under a prior mean μ carrying k pseudo-observations, the estimate is the Beta posterior mean
$$\widehat{\pi} = \frac{a + k\mu}{m + k}.$$
The prior mean at each level is the level above it: a cell is shrunk toward its source’s overall rate, a source toward the population’s rate, and the population toward the registry’s starting figure for that kind of fact. This is hierarchical shrinkage in the sense of Efron and Morris (1975). A cell with no observations answers with its source’s rate; a source with none answers with the population’s; a population with none answers with the registry prior.
Counts are discounted toward the present. With a decay factor λ per period, a count from t periods ago enters as λt of itself, giving an effective memory of 1/(1−λ) periods. At the default λ = 0.97 per week that is about 33 weeks.
Agreements are held in three separate channels by the independence of the judgment that produced them, and combined with a credit κb per channel:
$$a = \sum_{b}^{}\kappa_{b}\, a_{b},\quad\quad m = \sum_{b}^{}\kappa_{b}\, m_{b}.$$
The three bases are hard (independent ground truth, credit 1.0), dependent (a person confirming a value the record had already shown them, credit 0.25) and consensus (a source agreeing with the value the record settled on, credit 0.25).
Beside the reliability, the record learns each source’s PRECISION: the variance of its readings about its own systematic error, estimated from the residuals between its readings and an independent reference, on the same hierarchy and the same fading memory. The mean of those residuals is the source’s bias, and the two are held apart, because a biased source is precise and wrong at the same time and the two faults call for different responses. Only judgments in the hard channel contribute a residual, for the same reason only they carry full credit: a residual measured against the record’s own settled answer measures the record, not the source.
The geometry
With n items, the reliabilities (r1,…,rn) live in an n-dimensional unit cube. A binary record can occupy only its 2n corners: every item fully in or fully out. A fixed-weight record occupies one point inside the cube, chosen in advance. An evidence-weighted record holds a probability distribution over the cube, tied to the subject’s state θ, and the distribution moves as data arrive.
The same weighting logic repeats at every level of the record. A reading’s reliability is informed by its source. A source’s reliability at one kind of fact is informed by its record across all kinds. That structure is the hierarchy above, and it is what lets a subject with a sparse record benefit from what the source has shown elsewhere.
Proved: the record holds strictly more usable information
Theorem (total evidence). Let E be the evidence-weighted record and B = T(E) any binary record derived from it by an entry rule. For every prior and every loss, the Bayes risk attainable from E is no larger than the Bayes risk attainable from B, and strictly smaller whenever, on a set of positive probability, the Bayes action given B has strictly higher expected loss given E than the Bayes action given E.
Proof. Any decision rule dB that uses B defines a rule dE = dB ∘ T that uses E and has exactly the same risk. The Bayes rule for E minimizes risk over a set that contains every such dE, so its risk can be no larger. The strict case is immediate from the hypothesis. ▫
No distributional assumption is used. The result holds for any loss and any prior. Two boundaries on what it says, both respected throughout: it compares the BEST decisions available from each information set, so it does not by itself guarantee that a particular implemented estimator never does worse than a particular binary rule, which is what the no-regression result near the end is for; and the magnitude of the advantage is a separate question, which the measurements answer. This result is about direction, which is why it comes first.
It is the simplest case of Blackwell’s comparison of experiments (1951, 1953): a coarsened observation is never more informative than the observation it was coarsened from, for any decision problem at all. In information theory it is the data processing inequality (Cover and Thomas, 2006). In the philosophy of evidence it is Good’s principle of total evidence (1967). We claim no novelty for it. We claim that record systems are built as though it were false.
The result applies with full force to data that never enter a binary record. A family history left in free text, or a receipt photographed and never read, is an item with δi = 0, fixed by the schema before any analysis could judge its value.
Proved: our imperfection is bounded, and theirs is not
To measure the price of wrong weights, take the simplest setting where exact answers exist. Suppose each item is an unbiased reading of a single quantity with its own noise:
xi = θ + ei, 𝔼[ei] = 0, Var(ei) = σi2, independent.
Any nonnegative weights wi give the estimate ${\widehat{\theta}}_{w} = \sum_{i}^{}w_{i}x_{i}/\sum_{i}^{}w_{i}$. Among linear unbiased estimators the minimum-variance weights are proportional to 1/σi2 (Aitken, 1935). Write each item’s share of the total precision as $p_{i} = \sigma_{i}^{- 2}/\sum_{j}^{}\sigma_{j}^{- 2}$, and write ui = wiσi2 for the weight actually used relative to the correct weight. When ui is the same for every item, the weights are perfect.
$$\frac{Var\left( {\widehat{\theta}}_{w} \right)}{Var\left( {\widehat{\theta}}_{\text{best}} \right)} = \frac{\sum_{i}^{}p_{i}u_{i}^{2}}{\left( \sum_{i}^{}p_{i}u_{i} \right)^{2}} = 1 + {CV}_{p}^{2}(u).$$
Proof. $Var\left( {\widehat{\theta}}_{w} \right) = \sum_{i}^{}w_{i}^{2}\sigma_{i}^{2}/\left( \sum_{i}^{}w_{i} \right)^{2}$. Substitute wi = ui/σi2 and let $S = \sum_{j}^{}\sigma_{j}^{- 2}$. The numerator becomes $\sum_{i}^{}u_{i}^{2}\sigma_{i}^{- 2} = S\sum_{i}^{}p_{i}u_{i}^{2}$ and the denominator becomes $S^{2}\left( \sum_{i}^{}p_{i}u_{i} \right)^{2}$. The best variance is 1/S. Divide. ▫
The left side is the VARIANCE RATIO of the weighting used to the best weighting; the ratio minus one is the EXCESS variance, and the two terms are kept distinct throughout. The cost of misweighting depends only on how uneven the weighting errors are, measured across the information each item carries. Three consequences follow.
Coarse or imperfect weights have a ceiling. If every relative weight lies within a band [a, b] with b/a = κ, then
$$1 + {CV}_{p}^{2}(u) \leq \frac{(1 + \kappa)^{2}}{4\kappa}.$$
This is the Kantorovich inequality (1948); the proof is in the appendix.
Exclusion has a floor. If a binary rule rejects items carrying a share q of the total precision, the variance ratio is at least 1/(1−q), even when everything it keeps is weighted perfectly. Rejecting patient-reported values that carry 40 percent of the information costs at least 67 percent excess variance, before any other error.
Acceptance has no ceiling. A binary rule gives every accepted item the same weight, so its relative weights are proportional to σi2 and its κ equals the ratio of the largest to the smallest accepted variance. That ratio is set by the data, not by the design, and it grows without limit as noisier sources are accepted.
| Weight error band κ | Worst-case variance ratio | Equivalent share of the data thrown away |
|---|---|---|
| 1.5 | 1.04 | 4% |
| 2 | 1.13 | 11% |
| 3 | 1.33 | 25% |
| 4 | 1.56 | 36% |
| 9 (binary acceptance of a source with 9 times the variance) | 2.78 | 64% |
What κ is measured in, and how it is measured
The band in the figure is the one number here that comes from measurement rather than from a theorem, and both the quantity and the protocol matter.
κ is the spread of ui = wiσi2: each source’s weight multiplied by the variance of the source it weights. A source has two distinct properties a record could learn. Its RELIABILITY is the probability that a value from it is right. Its PRECISION is how far its readings scatter. The two are not proportional to one another, and a source can be right ninety-nine times in a hundred and still be the noisier of two. The ceiling is stated in precisions, so a spread of learned reliabilities, however small, certifies nothing about it. The record therefore learns each source’s precision directly, from the residuals between its readings and an independent reference, and the same runs measure what reliabilities would have given: weights taken from learned reliabilities sit at a κ near 5.5 in this configuration at every size, a ceiling near 90 percent, because reliability is simply not the quantity the ceiling is stated in.
The measurement protocol, in full. Four sources of true noise 2, 4, 6 and 9 report unbiased readings of true levels drawn uniformly on 60 to 240. Each run draws n readings per source, scores each against an exact reference, hands them to a fresh learner with no state from any other run, asks for the four precision weights, and takes that run’s κ as the largest ui over the smallest. 1,500 runs per size. The reported figure is the median run’s κ, with the 90th percentile beside it, because the ceiling is a per-run guarantee at that run’s κ and a median must not be read as a worst case.
| Judgments per source | Median κ | Ceiling at the median | 90th percentile κ | Ceiling at the 90th percentile |
|---|---|---|---|---|
| 25 | 1.74 | 7.8% | 2.49 | 22.3% |
| 100 | 1.32 | 1.9% | 1.56 | 5.1% |
| 1,000 | 1.09 | 0.20% | 1.16 | 0.53% |
| 10,000 | 1.03 | 0.019% | 1.05 | 0.053% |
Excess variance the learned weights can cost, measured on the shipped estimator, cold start per run. The two binary figures to set these against are 67 percent and 178 percent, at any size, forever.
Learning precision directly also does something an agreement rate cannot do: it separates a noisy source from a biased one. A reading far from the reference fails an agreement comparison either way, while a residual carries both a distance and a sign. The two faults call for opposite responses, and the sections below use the distinction.
Coarse weights beat binary rounding whenever their error band is narrower than the spread of reliabilities a binary rule accepts, or than the information it discards. A weighting could in principle be worse than a well-chosen binary rule if its weights were wildly wrong. The no-regression result below bounds that case.
Proved: a measured quantity is settled by combining
A record settles two different kinds of fact and they need two different estimators. Conflating them is the most common mistake in this area.
A POSTED fact has one true value and the question is which report is right. A diagnosis code, a counterparty, a date of service, a bank transaction. Two sources reporting 100.00 and 100.50 for one purchase are not reporting a transaction of 100.25; one is the amount and the other is a misreading, and averaging them produces a number that never happened. The estimator is a posterior over the reported candidates, and it is the one described in the preceding sections.
A MEASURED fact has one true value that nobody observes directly, and every report of it carries its own error. A serum potassium, a blood pressure, the market value of a property. Three automated valuations are three estimates of one number and the best estimate is none of the three. A rule that must return one of the reported values caps what it can achieve, because every reported value carries its reporter’s full noise.
The gap is easy to state and large. Four sources of equal noise measuring one quantity: a rule that selects one source’s reading in advance carries that source’s full noise, while the inverse-variance combination of all four carries half. Over 200,000 trials the measured errors are 0.9996 for selecting a source and 0.4986 for combining, against a theoretical 0.5000. A cleverer selection rule, such as taking the middle of the reported values, lands between those two figures and still above the combination, because any rule confined to the reported values forgoes the averaging that cancels independent noise; the figures above are for the source-selection rule, which is what a record that designates a “most trusted source” actually runs.
Three things the textbook form does not handle, and a record must.
Correlated sources enter the covariance, not the count. Inverse-variance weighting assumes independence exactly as a product of likelihoods does, and fails the same way. Two feeds derived from one upstream reporting 100.0, against one independent source reporting 106.0: treated as three witnesses the answer is 102.0 with a standard error of 0.577; with the family declared, the answer is 103.0 with a standard error of 0.707. The naive treatment is wrong twice, moving the answer toward the duplicated reading and claiming more precision than three readings from two witnesses can support. At full correlation the general solution reduces to the family rule already used for categorical facts, which is a sign the two estimators are one model rather than two.
A biased source is removed, not down-weighted. A source whose residuals have a direction is not noisy. It is wrong in a way that does not average out, and down-weighting only slows the damage: with enough readings it still moves the answer by its offset. The case that makes this matter most is the one where the biased source is also the most precise, which is common, because a well-calibrated instrument reading consistently high is precise and wrong at the same time. Two honest sources of variance 4 reporting 100, and one source of variance 0.01 reporting 130 with a measured offset of 30: weighting by precision alone returns 129.85, because the biased source carries almost all of the precision in the set. With the bias test the answer is 100.00 exactly. Learning precision without testing for bias is worse than learning nothing.
What a source reported is never replaced. The combination is a new number that sits beside the reported readings; it does not overwrite them, and no reported value is hidden from a reader on the strength of what the record believes. Where a source’s error fits a single constant the record offers its own estimate of the true value behind that reading, as a separate figure. Where the error is proportional rather than constant it offers nothing, because one subtraction would leave such a source wrong at both ends of its range and right in the middle, with the fault no longer visible to anyone reading the corrected number.
The interval is measured, not asserted. Three sources of noise 1, 2 and 5 measuring one quantity, over 100,000 trials: the stated 95 percent interval contains the truth in 94.993 percent of them. An interval that does not cover at its stated rate is a false claim, so coverage is measured rather than derived.
Telling the two faults apart
A source can be bad in two ways that call for opposite responses, and an agreement rate cannot distinguish them, because a reading far from the reference fails the comparison either way.
Two sources right equally often, 89.85 percent each, one landing close when it misses and one landing far: their learned variances are 0.038 and 5.28, a ratio of 140 in the weight each should carry in a combination. Nothing in their agreement rates separates them. Neither is biased, so neither should be set aside; one should simply count for far more than the other.
The direction of the residuals is what separates noise from bias, and the slope of residual against the REFERENCE level is what decides whether a correction would be sound. The regressor matters, and it is the reference, never the source’s own reading. A residual contains the source’s own noise, and so does the reading it came from, so the two share an error term: regressed on its own reading, even an honest source shows a positive slope of Var(noise)/Var(reading), and a constant-offset source, the one case subtraction honestly fixes, shows the same artifact and with enough readings is refused the correction it deserves. Regressed on the reference, an honest source and a constant-offset source both sit at zero in expectation, and a source reading proportionally high by a factor c has slope c − 1 exactly, which is the quantity the test is after. A reference carrying its own noise u pulls the slope toward − Var(u)/Var(reference), a small conservative bias when references are sound; the test’s refusals are therefore slightly more frequent than they need to be, never less.
Four sources, each with 300 readings of true levels spanning 60 to 240, honest noise 3, run on the shipped model:
| The source | Bias, in standard errors | Slope, in standard errors | The record’s conclusion |
|---|---|---|---|
| Honest | 0.0 | -0.9 | Unbiased; weight it by its precision |
| Reading a constant 3 high | 16.7 | -0.9 | Biased, one constant explains it; correctable, and the estimate is offered beside the reading |
| Reading ten percent high | 37.0 | 13.0 | Biased, and no constant explains it; correction refused |
| Reporting in the wrong unit | 40.5 | 14.6 | The same refusal, for the same reason |
| Every reading at one value | 1.3 | -17.3 | Readings carry no information about the level; correction refused |
The last row is the one to read twice. A source that reports the same value whatever the truth shows a slope of exactly − 1: its residual falls one for one as the true level rises, which is the signature of readings that do not track the level at all. Its mean residual can sit near zero, so a bias test alone would pass it; the slope is what convicts it, and its learned variance, which is the variance of the truth itself, is what strips its weight. The genuinely untestable case is different: a source compared only ever at one true level gives the slope nothing to vary against, and an untestable fit is reported as no fit, never as a fit, because the cost of the first is some accuracy and the cost of the second is a correction that hides an error. The record fails closed, and a product that sends residuals without the readings they came from gets no corrections at all.
Proved: what a value is worth depends on the rest of the record
A binary entry rule looks at one item at a time: this value, this source, this label. The reliability an item deserves depends on more. Under the model, the probability that item i means what it claims, given everything in the record, is
$$\Pr\left( z_{i} = 1 \mid E \right) = \frac{\pi_{i}\int f\left( x_{i} \mid \theta \right)\, p\left( \theta \mid E_{- i} \right)\, d\theta}{\pi_{i}\int f\left( x_{i} \mid \theta \right)\, p\left( \theta \mid E_{- i} \right)\, d\theta + \left( 1 - \pi_{i} \right)\int g\left( x_{i} \mid \theta \right)\, p\left( \theta \mid E_{- i} \right)\, d\theta}$$
where πi = π(si,ci;ϕ) and E−i is the rest of the record. The rest of the record enters through p(θ∣E−i). A value that agrees with everything else known earns more belief. A value that contradicts it earns less, in proportion to how strongly the rest of the record is established.
The proof is the same argument as the total-evidence result: per-item rules are a restricted class, and minimizing over a larger class cannot give a larger minimum. The coupling runs in both directions. Admitting a datum at its earned weight shifts the estimate of θ, which shifts the weight every other datum deserves.
A worked example: one glucose value
A patient’s record shows a hemoglobin A1c of 5.9 percent, corresponding to an estimated average glucose of 123 mg/dL (Nathan et al., 2008), and earlier fasting glucose values near 105 mg/dL. A new result arrives: glucose 132 mg/dL, labeled fasting. The American Diabetes Association uses 126 mg/dL as the fasting cutpoint for diabetes.
A binary record stores 132 as a fasting value above the cutpoint, and every rule that reads the record treats it that way, whatever else is known.
An evidence-weighted record computes the probability that the value really is a fasting value and the probability that the person’s true fasting glucose is at or above 126. For exposition we use a prior on true fasting glucose centered at 105 mg/dL with a standard deviation of 12, a true-fasting reading varying by 8 mg/dL around the true level, and a non-fasting reading running 35 mg/dL higher with a spread of 25. In practice each of these is estimated from outcome data.
| Context of the 132 mg/dL value | Prior reliability | Reliability given the whole record | Probability true fasting glucose is 126 or higher |
|---|---|---|---|
| Drawn in clinic, fasting confirmed by staff | 0.97 | 0.92 | 0.34 |
| Labeled fasting, not confirmed | 0.85 | 0.66 | 0.25 |
| Labeled fasting, but the patient told the guide she ate breakfast | 0.25 | 0.10 | 0.05 |
| Binary record, any context | 1 (by rule) | 1 (by rule) | treated as above the cutpoint |
The same number carries three different weights of evidence depending on its context and on the rest of the record. Even the best-documented reading supports only a one-in-three probability that the true fasting level is in the diabetic range, and the recommendation that follows is a repeat test, which is what clinical guidelines require before a diagnosis. The binary record has already recorded a diabetic-range fasting value, and every downstream rule, alert and research query inherits it.
The same shape in a financial record
A client’s file shows two years of deposits averaging $14,200 a month. A new statement line arrives: a deposit of $61,000, which a lender’s rule would treat as qualifying income.
| Context of the $61,000 deposit | Prior reliability | Reliability given the whole record | Probability it is recurring income |
|---|---|---|---|
| Matched to a closed statement and to an invoice | 0.97 | 0.91 | 0.67 |
| Reported by one aggregator, unmatched | 0.80 | 0.42 | 0.19 |
| Read from a photographed statement, no match | 0.60 | 0.17 | 0.06 |
| Binary ledger, any context | 1 (by rule) | 1 (by rule) | treated as income |
The structure is the same, and so is the failure the weighting avoids: a single large deposit that happens to be a transfer, an insurance settlement or a loan draw becomes qualifying income in a binary ledger, and the file that depends on it is wrong in a way nothing in the file reveals.
Proved: a conventional record grows confident and wrong, silently
Value is about how much a record can tell us. Reliability is about whether what it tells us is true, and in particular whether its stated confidence can be trusted. This is where the difference between binary and weighted records is largest, and where it grows with the size of the record.
Suppose a fraction ε of the values a binary rule accepts are corrupted, with an average offset Δ. The record estimates a quantity by treating every accepted value as true. Its estimate carries a bias b = εΔ that does not shrink with more data. Its standard error, $\sigma/\sqrt{n}$, does shrink. Its nominal 95 percent confidence interval therefore narrows around a point that is in the wrong place.
$$C(n) = \Phi\left( 1.96 - \frac{\sqrt{n}\, b}{\sigma} \right) - \Phi\left( - 1.96 - \frac{\sqrt{n}\, b}{\sigma} \right),$$
which falls toward zero as n grows. Coverage reaches 50 percent at about n1/2 = (1.96 σ/b)2 records.
Proof. The estimate is approximately normal with mean θ + b and standard deviation $\sigma/\sqrt{n}$. The interval contains θ exactly when the standardized estimate falls within 1.96 of $- \sqrt{n}\, b/\sigma$. As n grows that center moves away from zero without bound, and the covering probability goes to zero. Setting $\sqrt{n}\, b/\sigma = 1.96$ gives the stated half-coverage point. ▫
A record of this kind becomes more confident and less correct at the same time. At small sizes the bias hides inside wide intervals and the record looks sound. As records accumulate, the intervals shrink past the bias and the stated confidence becomes false. The larger and more successful the record, the further its stated confidence departs from the truth, and nothing in the record itself signals it.
Because n1/2 scales with 1/b2, halving the remaining bias quadruples the size a record can reach before its conclusions fail.
Proved: the record finds the biased source, and a conventional one cannot
A weight can be used two ways, and the difference decides this result.
Used as a multiplier, a weight does not cure a bias. A biased source keeps a positive weight, and at that weight it contributes its full offset to any weighted mean. If the weights are wi and one source carries bias Δ, the residual bias of the weighted estimate is $w_{\text{bad}}\Delta/\sum_{i}^{}w_{i}$, which is smaller than the unweighted figure and is not zero. Down-weighting a biased source is not enough, and a record that only down-weights goes over the cliff with the rest.
Used as a detector, the same learning is what tells you which source to remove. That is the use this record is built for, and it is the one a conventional record cannot imitate, because it never recorded which source said what and so has nothing to score.
The protocol, in full. Three sources report a quantity whose true value is 100: a clinic draw with noise 2 and no bias, a home meter with noise 6 and no bias, and an outside fax with noise 9 and a bias of +7. A run at n records is COLD: a fresh history of about n records, n/3 readings per source, handed to a learner with no state from any other run or any other size. Every reading in the history has been scored against a repeat test: it agrees when it lands within 5 mg/dL of the truth, and its residual feeds the precision learner. The same tolerance applies to every source, so the label never knows which source is which. 240 runs at each size, on the shipped estimator. “At 25 records” therefore means a learner that has seen about 8 scored readings per source, and nothing else.
The removal rule is a test on the mean residual, read in units of its own standard error: a source whose mean residual sits more than four standard errors from zero is reporting a systematic error, and averaging it in moves the answer however many readings arrive, so it is removed, and the survivors are weighted by learned precision. Four standard errors is a chosen false-positive rate; what is not chosen is the scale, because the statistic is standardized, and the test therefore tightens as evidence arrives. Nothing is told to the rule and nothing outside the record is consulted.
The sign and size of the mean residual are the whole of the rule, and for a reason. A rule that instead drops any source whose agreement rate falls below some fraction of the best source’s detects DISAGREEMENT with the reference, and a noisy honest source produces disagreement just as readily as a biased one. Against a precise biased source paired with a noisy honest one, an agreement-rate rule removes the honest source, keeps the biased one, and lands further from the truth than removing nothing at all; that adversarial configuration is part of the simulation suite. A mean residual carries a sign and an agreement rate does not.
Every method in the comparison can be built and run by a real system, and none of them is told anything it could not find out for itself.
| Method | What it is given at the start | Can it be built? |
|---|---|---|
| Accept every value | nothing; this is the conventional record | yes |
| Fixed weights, never revised | the registry’s starting figures | yes |
| Learned weights used only to reduce a source’s influence | learned reliabilities, applied as multipliers | yes |
| Holarc and Finarc: find the bad source, remove it, weight what remains | learned residuals, driving a removal decision | yes |
| Records | Accept every value | Fixed weights | Learned, influence only | Holarc and Finarc |
|---|---|---|---|---|
| 25 | 65% | 78% | 88% | 100% |
| 100 | 5% | 15% | 54% | 100% |
| 1,000 | 0% | 0% | 0% | 100% |
| 10,000 | 0% | 0% | 0% | 100% |
How often each method’s stated 95 percent interval contained the truth. This is coverage, which is what a confidence interval promises: a continuous estimate is never exactly right, and the question that can be answered is whether the interval a method states around it does the job it claims.
| Records | Accept every value | Fixed weights | Learned, influence only | Holarc and Finarc | Best linear weighting of the honest sources |
|---|---|---|---|---|---|
| 25 | 2.47 | 1.93 | 1.68 | 0.546 | 0.501 |
| 1,000 | 2.34 | 1.87 | 1.19 | 0.0757 | 0.0750 |
| 10,000 | 2.34 | 1.87 | 1.18 | 0.0244 | 0.0245 |
Median absolute error of the estimate, mg/dL. The last column is the two honest sources weighted by the inverse of their true variances, which Aitken proved is the minimum-variance linear unbiased weighting of those sources; it requires the true noises and so can be built by nobody.
Five results, in the order they matter.
Three methods promise 95 percent coverage and deliver zero. At one thousand records and beyond, every conventional method in the figure stated a 95 percent interval that contained the truth zero percent of the time. The intervals were not merely optimistic, they had stopped containing the answer entirely, and nothing in those records reported it. Ours covered in 2,158 of 2,160 runs across the nine sizes, never below 99.6 percent at any size, which is conservative against its own promise of 95.
The intervals are honest before the detection fires. At 25 and 50 records the bias test has nowhere near the evidence to convict anyone, and it convicts no one. The estimate stays covered anyway, because the precision weights already hold the noisy biased source to a small share of the answer and the interval is wide enough to carry what remains. The removal is what lets the error keep falling afterward; the weighting is what keeps the record honest while the evidence accumulates. A method that needed the detection to be early in order to be safe would be a different and worse method.
The record finds the biased source when the evidence supports it, and not before. The test removed exactly the biased source and nothing else in 45 percent of runs at 100 records, 99 percent at 250, and at least 99.6 percent at every size from 500 on. At 25 and 50 records it fired in none, and should not have: with 8 to 16 residuals of noise 9, a bias of 7 is about two standard errors from zero, which is evidence of nothing at a four-standard-error rule. A reader who checks the arithmetic will find the detection curve exactly where the standardized statistic puts it.
At ten thousand records we are 96, 77 and 48 times closer to the truth. Our median error is 0.0244 mg/dL. Accepting every value gives 2.34, which is 96 times further off. Fixed weights give 1.87, 77 times further. Learned weights used only to reduce a source’s influence give 1.18, 48 times further. These are ratios between methods measured in one simulation, never figures assembled from separate studies.
More data rescues ours and does not rescue theirs. Between twenty-five records and ten thousand our error falls by a factor of 22, from 0.546 to 0.0244, and it is still falling at the right-hand edge of the chart. The three conventional lines are flat: accepting every value sits at 2.47 and then 2.34, no improvement at all across a four hundred fold increase in data. A record carrying a bias cannot buy its way out with volume.
How much is still on the table
Approximately none, and the size of the remaining gap is measured rather than guessed at. A caution first about what the reference is: weighting the two honest sources by the inverse of their TRUE variances is the minimum-variance LINEAR UNBIASED weighting OF THOSE SOURCES (Aitken, 1935), compared here on median absolute error, where under this simulation’s normal errors the variance ordering and the median-error ordering agree. It is not the best possible estimator of the quantity: an estimator that modeled and corrected the bias rather than dropping the source could do better, and so could one exploiting the sources’ error patterns. It is a yardstick for the weighting step alone.
Against that yardstick, the shipped estimator’s median error at ten thousand records is 0.0244 and the yardstick’s is 0.0245, a ratio of 1.005. The gap is within half of one percent, which is within the resolution of the measurement. Nothing can be built that is told σi, because it is the true noise of a source, a property of the world and of no record; the yardstick is reported so the headroom is stated here, and there is approximately none of it against the class it defines.
The weighting step itself can be isolated on the same survivors. Weighting them by learned reliability gives 0.0301; weighting them by learned precision gives 0.0244, an improvement of 19 percent from using the quantity the theorem calls for. A reliability is a probability of being right, a precision is the inverse of a scatter, and the optimal weight for combining continuous readings is the second.
Down-weighting alone is not enough
The line for learned weights used only as multipliers holds its coverage longer than any conventional method, to 88 percent at 25 records and 54 percent at 100, and then goes over the same cliff. A record that learns reliabilities and does nothing with them except scale values will fail in exactly the way a conventional record fails, only later. The learning has to drive a decision. A system that got this wrong would look like ours and behave like theirs, which is why the failure is drawn in the figure.
The cliff theorem is untouched by any of this. It describes what happens to any estimator carrying a fixed bias, and it is why the three conventional lines in the figure end where they do.
Proved: one wild value cannot move the answer
The same contrast appears one value at a time. For a record that accepts a value and averages it in, the influence of that value on the result grows without limit as the value moves away from the truth (Hampel, 1974): one wild value, accepted, moves the answer as far as it likes. A binary exclusion threshold cuts influence to zero past a line and leaves it at full strength just inside the line.
Under a mixture model, a value’s influence on the estimate is multiplied by its posterior reliability, which falls smoothly as the value becomes less consistent with the rest of the record. When the model for bad values is broad and does not track the true state, that influence is bounded and fades out instead of stopping at a line. Huber (1964) and Hampel (1974) established this family of estimators as the standard for resistance to contamination.
In the shipped combiner, the equivalent protection is structural: a lone source reporting a value nothing else supports produces a settled fact whose confidence is that source’s own weight, and a value no source supports is reported as the explicit probability that the truth is something nobody reported. The wild value does not move the answer; it lowers the confidence in it.
What outcome history is worth, and where it is decisive
The results so far are about weighting. This one is about learning, measured on the model that ships. It also answers the question a skeptical reader should ask of any such claim: where does this not help, and how would we know?
Three regimes, each run over six seeds with 8,000 trials per seed. The truth is one value; a source that is wrong reports one of twenty alternatives, or its own systematic error. Sources differ in reliability and in how often they are present at all. A fourth column is added that no product can use and every fair comparison needs: the error when every source is weighted by its true reliability. One caution about that column: it is an oracle input to THIS combiner, the best any weighting handed to this vote can achieve, and claims about “the best any weighting can do” below mean exactly that. It is not an optimum over all conceivable inference methods.
| Regime | Starting figures | Learned | Learned, pair declared one witness | Oracle weights |
|---|---|---|---|---|
| A source that silently broke | 51.5% | 18.4% | not applicable | 18.3% |
| A correlated pair outvoting a good source | 23.4% | 22.0% (15.9 to 23.6) | 16.1% (15.8 to 16.3) | 23.4% |
| Independent mistakes | 1.63% | 1.61% | not applicable | 1.63% |
Settled facts that are wrong. Ranges are across seeds where the spread matters. Oracle weights means every source weighted by its true reliability, which no system can know and which bounds what any weighting handed to this combiner can achieve.
A source that silently broke. A feed that was reliable and now reports a systematically wrong value, still present on 95 percent of facts. Learning takes the error from 51.5 percent to 18.4 percent, and oracle knowledge of every source’s true reliability would take it to 18.3. The learned weights capture 99.7 percent of everything available to be won. This is where outcome history earns its place, and it is a case a conventional record cannot address at all, because it kept no record of which source was right before and so has nothing to learn from.
A correlated pair. Two feeds off one upstream, both reporting the same wrong value, against a good source that is often absent. Learning marks the pair down correctly, to about 0.62 from a starting 0.80, and still leaves the outcome to chance. The combiner’s decision between the pair and the good source sits on a threshold near a good-source weight of 0.982, and which side a run lands on is noise: across six seeds the error ranges from 15.9 to 23.6 percent. Declaring the two sources one witness gives 16.1 percent on every run, with a range under half a point. Declaring the family using the starting figures and no learning at all does equally well.
Oracle weights settle this regime at 23.4 percent, which is the starting figure to two decimal places. No weighting of any kind helps here, including the best one that exists. The declaration reaches 16.1 percent, so it achieves something that perfect knowledge of every source’s reliability cannot. The declaration is a separate mechanism from weighting, and a conventional record has no field, no concept and no remedy for what it addresses: there is no place in such a system to say that two institutions are one witness.
The conclusion is specific: in this regime the declaration does the work and the learning adds nothing. The result runs against the intuition that learning should help everywhere, and it is reported the way the model behaves.
A note on what the declaration is and is not. Copying between sources can also be INFERRED from data, by the correlation of errors: sources that share false values beyond chance are probably not independent, and the truth-discovery literature does exactly this (Dong, Berti-Équille and Srivastava, 2009). The declaration differs on three points that matter to a record of record: it is certain where inference is probabilistic, it protects the first contested fact rather than requiring a history of shared errors to accumulate, and it is auditable, because it states an administrative fact about institutions that a person attested, rather than a statistical conclusion that shifts with the data. A deployment that wants both can run inferred-copying detection inside NERD as a flag for families nobody declared; the record’s own counting rests on the declaration.
Independent mistakes. Four sources making honest, uncorrelated errors. The starting figures give 1.63 percent, learned weights 1.61, and weighting every source by its true reliability gives 1.63. All three are the same number inside the run-to-run noise.
Learning is not failing here; the problem has no room in it. When independent sources disagree, the outcome is decided by who agrees with whom, and no assignment of weights, including the oracle’s, changes many votes. The ceiling and the floor are the same figure, and every method is already standing on it.
Two things follow. A reader can take the 64 percent and 31 percent figures above more seriously, because the same measurement apparatus reports no gain where none is available instead of manufacturing one. And the no-regression guarantee below, which bounds how far the record can fall behind a conventional one, is visible here as a measurement: handed a problem where a conventional record is already at the reference figure, the record matches it.
The residual 16 percent is that floor. Those are the facts the pair alone witnessed, where nothing contradicts them. The record’s obligation there is to report truthfully rather than to be right, which is what the confidence-if-independent figure below is for.
Why memory must fade
A source that was reliable and has drifted should not keep coasting on old credit. A feed reliable at 97 percent for a year, then silently dropping to 45 percent, with twelve judgments a week:
| Time since the break | Fading memory, λ = 0.97 per week | Every judgment counted forever |
|---|---|---|
| 6 months | 0.669 | 0.775 |
| 1 year | 0.532 | 0.684 |
| 2 years | 0.469 | 0.610 |
| 3 years | 0.436 | 0.570 |
The fading memory converges to the true 0.45. Without it the weight falls only as fast as the old history dilutes, and after three years it still reads 0.57. Neither is fast: a broken feed carries a materially wrong weight for about a year either way. The decay sets how long, and the right value of λ is a deployment’s judgment about how fast its sources change.
The trade bought here is named rather than hidden. A fixed decay caps the effective evidence behind a weight at about 1/(1−λ) periods of judgments, so a decayed weight never becomes arbitrarily certain however long the history runs: for a STABLE source, some estimation noise is kept forever that counting every judgment would have eliminated. That is the price of tracking, and the weight’s own account of itself reports the effective evidence behind it, so a reader can see how much is paid. The decayed estimate tracks a changing truth and is not a consistent estimator of a fixed one; a record that preferred consistency over tracking would be the record still trusting the broken feed three years on.
Why the independence of a judgment must be recorded
A record that shows a reader its current belief and then counts that reader’s agreement as evidence is grading its own work. The same is true of a source that reports a value after the record has settled on one.
A source whose true reliability is 0.70, which the record keeps agreeing with, so it “agrees” 97 percent of the time regardless of its real quality:
| Policy | Learned weight after 400 judgments |
|---|---|
| Agreement counted as independent truth | 0.952 |
| Agreement discounted to a quarter, no ceiling | 0.934 |
| Agreement recorded as agreement, with a ceiling | 0.773 |
Why no fixed discount is enough. Discount self-confirming agreement by any fixed credit κ > 0 and the estimate is π̂ = (κa+kμ)/(κm+k). Divide numerator and denominator by κm and let m → ∞: the estimate converges to a/m, the raw agreement rate, whatever κ is. ▫ A source agreeing with the record 97 percent of the time is therefore carried toward 0.97 by any fixed discount; the discount sets the speed of the climb and never its destination.
This is why the model caps rather than only discounts. A weight whose evidence is mostly self-confirming is not allowed to rise above what the independent channels support. The ceiling is one-sided: agreement that disagrees still pulls a weight down, because a source contradicting the settled value is informative either way. In the same setting, a source agreeing only 100 times in 400 falls to 0.27.
The one-sidedness has a cost, and it is stated in The boundary below: when the record’s settled answer is wrong and a minority source is right, the minority’s disagreement pulls its weight down for being right. The hard channel exists to repair exactly this, because an independent outcome restores the weight the consensus took away, and the repair is only as fast as independent outcomes arrive.
The credit per channel is a tuning number, and what makes the ceiling possible at all is that the basis of each judgment was recorded at the time. A system that merges its labels into one count can never apply this ceiling later, because the information needed to apply it is gone.
Proved: a contested fact is settled with its assumptions visible
When sources report different values for one fact, the combiner returns a posterior over the reported values and over the explicit hypothesis that the truth is a value nobody reported. With m reported values, a budget of A alternatives for the unreported case, and source weights ws:
$$\Pr(T = v)\mspace{6mu} \propto \mspace{6mu}\frac{1}{m + A}\prod_{\text{families}}^{}\ell_{f}(v),\quad\quad\Pr\left( T \notin \text{reported} \right)\mspace{6mu} \propto \mspace{6mu}\frac{A}{m + A}\prod_{\text{families}}^{}\ell_{f}(\varnothing),$$
where each family of correlated sources contributes one factor at its most reliable member’s weight. The settled fact reports the most probable value, its posterior, every rejected value with its own posterior, which sources were counted as one witness, and the confidence the answer would have carried had those sources been treated as independent.
That last figure is what the independence assumption would have cost. Two feeds off one upstream, both reporting the same wrong value at weight 0.62, read as 0.98 confident when treated as independent and as 0.62 when declared one witness. The answer is the same either way; the confidence is not, and the confidence is the number that goes in front of a person. A conventional system has no way to express it.
Proved: the weights converge, and the convergence is measurable
The result. Under the usual regularity conditions for maximum likelihood and Bayesian estimation (van der Vaart, 1998), with outcome-linked judgments accumulating for a source whose behavior is stable, the learned reliability converges to the true reliability and the learned precision to the true precision, at the standard $1/\sqrt{m}$ rate in the effective count m, up to the floor the fading memory imposes. The measured κ above, whose median falls from 1.74 at 25 judgments per source to 1.03 at ten thousand, is this result observed on the shipped estimator, in the quantity the Kantorovich ceiling is stated in.
Three features turn the property into a living system. Learning happens per cell, so a source good at one kind of fact and bad at another is described correctly, never averaged. Weights change without editing the record, because every event keeps its source and the weight in force is derived at read time. And a new source with no track record starts at a conservative figure and earns more as outcomes accumulate, where a binary system has two options for a new source: trust it fully or ignore it.
One qualification, and the regimes measured it. Convergence of the weights is not the same as improvement in the answer. Against independent noise the weights converge to the true reliabilities and the error does not move, because weighting by the true reliabilities does not move it either. Convergence is worth exactly what the problem has in it to win, which in that regime is nothing and in the broken-source regime is a 64 percent reduction in wrong facts.
Proved: the record cannot fall meaningfully behind a conventional one
The result. Binary rules are special cases of weightings, so any score a binary rule can earn, some weighting in the family earns too. When the weighting in force is chosen or validated against outcome-linked data held out from training, the uniform-convergence bound for selection by validated loss (Shalev-Shwartz and Ben-David, 2014) applies: with probability at least 1 − δ, the deployed weighting’s true loss exceeds the best binary rule’s by at most $O\left( \sqrt{\log(N/\delta)/m} \right)$, where m is the validation count and N indexes the rules compared. The gap shrinks as validation outcomes accumulate, and it is the price of not knowing the best rule in advance, paid by every learning system and bounded for this one.
It is applied operationally in three ways. Weights start at the conventional figure for each kind of fact and move away only as validated outcomes support the move. Calibration is measured and reported. And a source whose calibration drifts is flagged for a person before its weight moves further.
The guarantee is also visible as a measurement. In the independent-mistakes regime a conventional record is already at the reference figure, at 1.63 percent, and the record returns 1.61. Handed a problem with nothing in it to win, it wins nothing and loses nothing. That is the behavior the bound promises, observed.
Scope: what each result assumes, and what would falsify it
Each result is a theorem or a measurement, and each holds under conditions. They are set out here because a claim whose conditions are stated can be checked, and one whose conditions are hidden cannot. Nothing in this table withdraws a result above; each line says where the result stops applying.
| Result | Assumes | Breaks when |
|---|---|---|
| Total evidence | The binary record is a function of the weighted one; best-achievable decisions are compared | The binary record contains something the weighted one does not, which cannot happen by construction |
| Rounding-cost identity, Kantorovich ceiling | Unbiased independent readings of one quantity | Readings are biased, which the cliff section handles, or correlated |
| Coupling | A model linking items through the subject’s state | Items are independent given the source, in which case the gain is zero |
| Binary cliff | A fixed bias that does not shrink with the sample | The bias is zero, or shrinks with the sample |
| Finding the biased source | An independent reference to score residuals against, and enough of them for the standardized test | No reference exists, or the bias is too small for the evidence, so the test stays its hand and the precision weighting carries the protection |
| The three regimes | The simulated generating process; sources present and wrong at the stated rates | Real sources behave differently, which only a deployment can show |
| Correlated-witness floor | The correlated sources are the only witnesses | A third independent source is present |
| Self-confirmation | The record shows its belief before the judgment is made | Judgments are collected blind, which is what the hard channel is for |
| Learning result | The model is identified, the source is stable, and outcomes are linked | Outcomes are absent, so the weights stay at their starting figures; or the source drifts, which decay tracks at the stated cost |
| No-regression | Validation data are outcome-linked and separate from training | Validation is reused for selection |
What was run and is not reported. One further method was simulated and is left out of every figure and table: an analyst told from outside which source is biased, who then averages the survivors equally. It reaches 0.036 against our 0.0244. It is excluded as an incoherent baseline, and for no other reason: it has perfect knowledge of which source to drop and no knowledge at all of how to weight what remains, so beating it demonstrates nothing about either method. The only comparison this paper makes is against methods that can be built, and the only references it reports are labeled as references.
Two conditions apply to everything here. Detectability: a corrupted value indistinguishable from a true one, with no contextual signal and no conflict with the rest of the record, cannot be detected by any method. Weighting cannot recover information that was never observed; it can only avoid discarding information that was. Outcomes are needed to learn: weights learned without outcome data are starting figures, and every learning result strengthens only as outcome-linked data grow. Neither record has been calibrated against real outcomes, because neither has yet held a real record.
How this relates to classical methods
| Method | What it does with a doubtful datum | Uses the rest of the record | Learns reliability from outcomes | Improves as data accumulate |
|---|---|---|---|---|
| Conventional relational record | Accepts as true or rejects at entry | No | No | No |
| Complete-case analysis | Drops the whole case if any value is missing | No | No | No; biased unless data are missing completely at random (Little and Rubin, 2019) |
| Outlier deletion and quality flags | Drops values past a threshold | Rarely | No | No |
| Multiple imputation (Rubin, 1987) | Fills missing values from a model | Yes, for missing values | No | Within a study |
| Measurement-error models (Carroll et al., 2006) | Corrects for known error variance | Yes | Only with validation substudies | Within a study |
| Inverse-variance weighting | Weights by known precision | No | No | No |
| Resistant estimation (Huber, 1964) | Reduces the influence of discordant values | Yes | No | Within a study |
| Power priors (Ibrahim and Chen, 2000) | Discounts an entire outside dataset | No | Partly | Across studies |
| Latent-class source reliability (Dawid and Skene, 1979) | Learns each source’s error rate | Yes | Yes | Yes |
| Uncertain databases (Trio: Widom, 2005) | Stores alternatives with confidences and lineage in a relational model | Yes, through lineage | No | No |
| Truth discovery with copying detection (Dong et al., 2009) | Learns source accuracy and infers copying from shared errors, per integration task | Yes | Yes, within the task | Within the task |
| Evidence-Weighted Record | Keeps every datum with provenance; weights each by its learned, context-dependent reliability; records the independence of every judgment behind that weight; counts correlated sources once | Yes, jointly | Yes, per source and per kind of fact, with a fading memory | Yes, continuously, without editing stored data |
Each of these solves part of the problem, and the two database rows deserve the direct comparison. Trio showed that a relational system can carry alternatives, confidences and lineage; it does not learn source reliability from outcomes, fade its memory, or record the independence of the judgments behind a confidence. The truth-discovery line learns source accuracy and infers copying, within one integration task over a snapshot of data; it is the closest relative of the trust model, as Dawid and Skene is of the reliability estimator. What is added here is the combination as a PERMANENT RECORD OF RECORD: every datum kept with provenance for a lifetime, reliability learned per source and per kind of fact with a fading memory, the independence of every judgment recorded at the time it was made so self-confirmation can be capped later, correlated sources declared and counted once from the first fact rather than inferred after the errors accumulate, and the weighting revisable forever without editing stored data.
Naming the category
We propose two terms.
An Evidence-Weighted Record is a record that keeps every datum with its source and context and assigns each a reliability weight learned from outcomes and revised as evidence accumulates, so that analyses use each datum in proportion to the evidence it provides. A search found no prior use of the term for a record or database.
The binary cliff is the failure described above: as a record that accepts some bad data as true grows, its confidence intervals narrow around a biased answer until they no longer contain the truth. It is a property of how conventional records handle evidence, and the name gives clinicians, lenders, researchers and investors a way to ask whether a given record is exposed to it.
Conclusion: how much better, and why
The short version
Four sentences, each backed by a theorem or a measurement.
- For every question, every scoring rule and every prior, the best decisions available from an evidence-weighted record are at least as good as the best available from a binary one, and strictly better whenever the discarded information would have changed the decision. That is a theorem, with no assumptions to argue about.
- On a record containing one biased source, every conventional method stated a 95 percent interval that contained the truth 0 percent of the time at a thousand records and beyond. Ours contained it in at least 99.6 percent of runs at every size, 2,158 of 2,160 in all, including the sizes where the bias is not yet detectable.
- Its answer was 96 times closer to the truth than accepting every value, 77 times closer than fixed weights, and 48 times closer than learning weights and using them only to reduce a source’s influence. All four methods can be built; the comparison is between things that exist, cold start per run, under a protocol stated in full.
- The one reference above ours is the minimum-variance linear unbiased weighting of the surviving sources with their true noises known, which can be built by nobody. The shipped estimator sits within half of one percent of it.
The places where this record buys nothing are set out under The boundary below, in the same plain terms.
The claim
An evidence-weighted record is better than a binary one in a sense that needs no qualification and no simulation: for every question it will ever be asked, under every way of scoring the answer, and under every prior, the best decisions available from it are at least as good as the best available from a binary record built from the same data, and strictly better whenever the information the binary record discarded would have changed the decision. That is the total-evidence result, and it is a theorem. There is no distributional assumption behind it, no tuning, and no case in which a binary record built from the same data supports better decisions.
The theorem compares what each record makes possible. That the shipped estimator realizes the advantage, and does not fall behind where there is no advantage to realize, is what the measurements and the no-regression bound are for. Everything else is about magnitude. The theorem settles the direction.
How much better
Three questions have numbers attached. The first is a measured ceiling, the second and third are measured on the shipped model, cold start per run.
How much does it cost to be wrong about a weight? Little, and the ceiling is exact per run. The worst case for weights spread over a band of ratio κ is a variance ratio of (1+κ)2/4κ, an excess of that minus one. Measured in the bound’s own terms, the median run sits at these ceilings:
| Weights | Ceiling, median run | Ceiling, 9 runs in 10 under |
|---|---|---|
| Learned, 25 judgments per source | 7.8% | 22.3% |
| Learned, 100 judgments per source | 1.9% | 5.1% |
| Learned, 1,000 judgments per source | 0.20% | 0.53% |
| Learned, 10,000 judgments per source | 0.019% | 0.053% |
| Binary: exclude 40% of the information | 66.7%, at any size | 66.7% |
| Binary: accept a source nine times noisier | 177.8%, at any size | 177.8% |
At 25 judgments per source, which is a thin record, the median run’s ceiling is already about 9 times smaller than the penalty for excluding weak sources and 23 times smaller than the penalty for accepting noisy ones. At 1,000 judgments it is about 330 and 880 times smaller, and at 10,000 about 3,500 and 9,300. This holds without any appeal to learning working well: the ceiling only requires the weights to be roughly right, and the measurement reports how roughly, with the spread across runs beside the median.
How many facts does the record get right that a conventional one gets wrong? It depends entirely on what is going wrong, and the answer has three parts.
| What is going wrong | Conventional record | This record | Fewer facts settled wrongly |
|---|---|---|---|
| A source quietly breaks and keeps reporting | 51.5% wrong | 18.4% | 64%, which is 99.7% of all that was available |
| Two sources are secretly one source | 23.4% wrong | 16.1% | 31%, below what oracle weights reach |
| Ordinary independent mistakes | 1.63% wrong | 1.61% | nothing, because oracle weights also give 1.63% |
How much of a biased source can the record undo? All of it, once the evidence supports the removal, and the record stays honest while the evidence accumulates. The bias test identified the biased source and nothing else in 45 percent of runs at 100 records, 99 percent at 250, and at least 99.6 percent from 500 on, exactly where the standardized statistic puts those rates; below that it correctly convicted no one, and the precision weighting held coverage at or above 99.6 percent anyway. Every conventional method held 0 percent of its stated intervals from a few hundred records on while still reporting 95 percent confidence, and every one of them was still narrowing around the wrong answer as the record grew.
| At 10,000 records | Share of stated intervals that held | Median error | How much further off than ours |
|---|---|---|---|
| Conventional: accept every value | 0% | 2.34 mg/dL | 96 times |
| Fixed weights, never revised | 0% | 1.87 | 77 times |
| Learned weights, influence only | 0% | 1.18 | 48 times |
| Holarc and Finarc: find it, drop it, weight the rest | 100% | 0.0244 | the benchmark |
Every row in that table can be built and run. That is the condition that makes the comparison mean something, and it is why no oracle, expert, or outside hint appears anywhere in it. The record learns each source’s precision directly because a reliability is not a precision, and on the same survivors, weighting by learned precision instead of learned reliability is worth 19 percent on its own.
How badly can a record fool itself? A source whose true reliability is 0.70, which the record keeps showing its own answer to, reads as 0.952 reliable if its agreement is counted as evidence. That is an overstatement of 0.25 in absolute reliability, 36 percent in relative terms, on no outside evidence whatever. Discounting that agreement to a quarter leaves it at 0.934, because no fixed discount bounds the limit, as proved above. Holding it to what independent evidence supports puts it at 0.773.
Why
Three mechanisms produce those numbers. None of the three is a place where a conventional record does the same job worse. They are jobs it has no way to attempt, because the information they need was thrown away at entry.
It keeps what the others round off. A conventional record stores a value. This one stores a value, its source, the method, the original document, the moment, and a weight. The total-evidence result applies only because that information still exists. A record that rounded reliability to zero or one at entry cannot recover it later at any price, which is why the gap grows rather than closing as both records accumulate data.
It learns which sources deserve trust, per kind of fact, with a memory that fades. This is what finds the quietly broken source, and the 64 percent reduction is the whole of its value. A conventional record cannot do this because it never recorded which source said what, so it has nothing to score. Note what the fading memory is for: without it, a source with ten thousand good observations takes years to be marked down, and the simulation shows a broken feed still reading more than a quarter too high three years after it broke.
It records where each judgment came from, and which sources share an upstream. We have found no record system, in products or in the research literature, that stores with each judgment whether it was independent of the record’s own output and uses that stored basis to cap what self-confirming agreement can prove; the closest published work learns source accuracy and infers copying within an integration task, and the comparison with it is drawn above. This mechanism is what keeps the first two from grading their own work. A record that counts its own agreement as confirmation climbs toward certainty on nothing, invisibly. A record that counts two feeds off one exchange as two witnesses reports 0.98 confidence where 0.62 is warranted. Both failures are silent in every conventional system, and both are structural: the fix requires having recorded, at the time, something a conventional record had no field for.
The boundary
A claim is only as good as its stated limits, so here are ours, stated as plainly as the claims.
Against ordinary independent mistakes, learned weights buy nothing, and neither does anything else: weighting every source by its true reliability gives the same figure. The outcome of a vote among independent sources is decided by who agrees with whom. A record that promised a general accuracy improvement from learned weights would be overselling them, and this is the regime that shows why.
A weight used only as a multiplier does not cure a bias. A record that learns reliabilities and does nothing with them but scale values goes over the same cliff, later. The learning has to drive a decision, which in the simulation is the removal of the source it identifies, and the removal fires only when the evidence supports it: at small sizes the test correctly stays its hand and the precision weighting carries the protection. Where no independent reference exists to score sources against, there are no hard labels, the weights stay near their starting figures, and none of this applies.
Against correlated sources, learning alone is a coin toss. What works is the declaration, which is an attested administrative fact about institutions; copying can also be inferred from shared errors, as the truth-discovery literature does, and the trade between the two is stated above. When correlated sources are the only witnesses to a fact, nothing recovers it. The record’s obligation there is to report truthfully, which is what the confidence-if-independent figure exists for.
The self-confirmation ceiling is one-sided, and the open side has a cost: when the record’s settled answer is wrong, a correct minority source’s disagreement pulls its weight down for being right. The hard channel repairs this as independent outcomes arrive, and no faster. A deployment whose sources include a lone truthful dissenter among correlated majorities should expect the repair to lag the harm, and that case is part of the adversarial suite rather than a discovery waiting to be made.
The fading memory trades consistency for tracking: a decayed weight keeps a floor of estimation noise forever on a stable source, as the price of converging on a drifted one. The effective evidence behind every weight is reported so the price is visible.
None of this is calibrated against real outcomes, because neither record has yet held a real one.
What it comes to
A binary record rounds every datum’s reliability to zero or one before anyone knows what question the record will face, and the rounding is irreversible. The theorem says the information lost that way can only hurt the decisions built on it. The measured ceiling says that carrying imperfect learned weights instead costs orders of magnitude less than the rounding did. The measurements say the learning is worth 64 percent of the wrong answers when a source breaks and 31 percent when sources are secretly one source, and nothing at all when the errors are honest and independent.
These are theorems and measurements, and anyone can check them: the protocol behind every simulated figure is stated in full, and the measurements run against the same trust model the records run, which is the only reason to believe a number in a paper matches a number in a system.
Appendix: proof of the ceiling on imperfect weights
Let ui ∈ [a, b] with 0 < a ≤ b, and let $m = \sum_{i}^{}p_{i}u_{i}$. Because (ui−a)(b−ui) ≥ 0 for each i, we have ui2 ≤ (a+b)ui − ab. Multiply by pi and sum:
$$\sum_{i}^{}p_{i}u_{i}^{2} \leq (a + b)m - ab.$$
Dividing by m2 gives a ratio of at most h(m) = (a+b)/m − ab/m2. Setting h′(m) = − (a+b)/m2 + 2ab/m3 = 0 gives m = 2ab/(a+b), which lies in [a, b], and
$$h\left( \frac{2ab}{a + b} \right) = \frac{(a + b)^{2}}{2ab} - \frac{(a + b)^{2}}{4ab} = \frac{(a + b)^{2}}{4ab} = \frac{(1 + \kappa)^{2}}{4\kappa},\quad\quad\kappa = b/a.$$
The bound is attained when the precision is split between items with relative weight a and items with relative weight b in the proportions that make m = 2ab/(a+b). ▫
References
- Aitken AC. On least squares and linear combination of observations. Proceedings of the Royal Society of Edinburgh. 1935;55:42–48.
- Blackwell D. Comparison of experiments. Proceedings of the Second Berkeley Symposium on Mathematical Statistics and Probability. 1951:93–102.
- Blackwell D. Equivalent comparisons of experiments. Annals of Mathematical Statistics. 1953;24(2):265–272.
- Carroll RJ, Ruppert D, Stefanski LA, Crainiceanu CM. Measurement Error in Nonlinear Models: A Modern Perspective. 2nd ed. Chapman & Hall/CRC; 2006.
- Cover TM, Thomas JA. Elements of Information Theory. 2nd ed. Wiley; 2006.
- Dawid AP, Skene AM. Maximum likelihood estimation of observer error-rates using the EM algorithm. Journal of the Royal Statistical Society, Series C. 1979;28(1):20–28.
- Dempster AP, Laird NM, Rubin DB. Maximum likelihood from incomplete data via the EM algorithm. Journal of the Royal Statistical Society, Series B. 1977;39(1):1–38.
- Dong XL, Berti-Équille L, Srivastava D. Integrating conflicting data: the role of source dependence. Proceedings of the VLDB Endowment. 2009;2(1):550–561.
- Efron B, Morris C. Data analysis using Stein’s estimator and its generalizations. Journal of the American Statistical Association. 1975;70(350):311–319.
- Good IJ. On the principle of total evidence. British Journal for the Philosophy of Science. 1967;17(4):319–321.
- Hampel FR. The influence curve and its role in robust estimation. Journal of the American Statistical Association. 1974;69(346):383–393.
- Huber PJ. Robust estimation of a location parameter. Annals of Mathematical Statistics. 1964;35(1):73–101.
- Ibrahim JG, Chen M-H. Power prior distributions for regression models. Statistical Science. 2000;15(1):46–60.
- Kantorovich LV. Functional analysis and applied mathematics. Uspekhi Matematicheskikh Nauk. 1948;3(6):89–185.
- Little RJA, Rubin DB. Statistical Analysis with Missing Data. 3rd ed. Wiley; 2019.
- Nathan DM, Kuenen J, Borg R, Zheng H, Schoenfeld D, Heine RJ. Translating the A1C assay into estimated average glucose values. Diabetes Care. 2008;31(8):1473–1478.
- Proakis JG, Salehi M. Digital Communications. 5th ed. McGraw-Hill; 2008.
- Rubin DB. Multiple Imputation for Nonresponse in Surveys. Wiley; 1987.
- Shalev-Shwartz S, Ben-David S. Understanding Machine Learning: From Theory to Algorithms. Cambridge University Press; 2014.
- van der Vaart AW. Asymptotic Statistics. Cambridge University Press; 1998.
- Widom J. Trio: a system for integrated management of data, accuracy, and lineage. Proceedings of the Second Biennial Conference on Innovative Data Systems Research (CIDR). 2005:262–276.
- Willemink MJ, Persson M, Pourmorteza A, Pelc NJ, Fleischmann D. Photon-counting CT: technical principles and clinical prospects. Radiology. 2018;289(2):293–312.
- American Diabetes Association Professional Practice Committee. Diagnosis and classification of diabetes. Standards of Care in Diabetes, published annually in Diabetes Care, Supplement 1.
- Stoutenburgh G. The Evidence-Weighted Record. CS Quantum Suite White Paper Series, Paper 1, Rev 6. October 7, 2026.
- Stoutenburgh G. Markets and Applications. CS Quantum Suite White Paper Series, Paper 3, Rev 6. October 7, 2026.
- Stoutenburgh G. Data Security in the Quantum Era. CS Quantum Suite White Paper Series, Paper 4, Rev 5. October 7, 2026.
Every simulated figure and every measured number in this paper is produced by running the same trust model the records run, under the protocols stated beside each result: cold start per run, fixed seeds, counts of runs and sizes as given. Simulations are simulations: they illustrate the theorems and check them numerically, and they do not measure performance on real records. Neither record has held a real patient record or a real financial account.
© 2026 VitaNexus Holdings, Inc. All rights reserved. Holarc, Finarc and NERD, including their software, designs and documentation, are the property of VitaNexus Holdings, Inc. Confidential.
The white papers · Paper 3 of 4 · Rev 6 · October 7, 2026
Markets and Applications
Where the Evidence-Weighted Record Sells, and What That Is Worth.
The full paper, exactly as published. By Greg Stoutenburgh; founders and inventors Greg Stoutenburgh and Faiz Chowdhury. Also available as PDF and Word on request.
Summary
The record described in Paper 1 is a foundation, and a foundation on its own sells to nobody. What sells is the software people use every day, built so that each piece writes to the record and reads from it, and so that the record gets better the more of them are in use.
Seven pieces on one foundation. A lifetime health record and a lifetime financial record. A patient’s app, a clinic system, an imaging platform and an evidence engine. Beneath all of them one sign-in, one encryption system, one gateway to outside AI, one record engine, and tamper-evident logs. Together they are the CS Quantum Suite.
The two records are the single repository for everything in the suite. No application keeps data of its own. The same repository is open to other companies under license: a licensee runs its own data in a partition of the record, sees only its own data, and gets the storage, the integrity chain, the trust model and the security of the whole system without building any of it.
Counting only the segments the suite sells into directly, published estimates put addressable spending at about $107 billion to $128 billion a year worldwide and about $30 billion to $36 billion a year in the United States. Three adjacent segments sit beside those numbers and overlap them in part.
The suite has no customers and no revenue. It runs on invented data and no real patient record or financial account has entered it. What follows is what each piece is, who buys it, what it sells into, how it earns, and what remains before any of it can be sold.
The papers in this series
- Paper 1, The Evidence-Weighted Record. What the record is, how the database works, and how a weight is learned.
- Paper 2, The Mathematics. The proofs, and the measurements of what learning from outcomes is worth.
- Paper 3, this document. The products, the markets, the business model, and the gap to revenue.
- Paper 4, Data Security in the Quantum Era. How the design keeps stolen data unusable.
The seven pieces
| Piece | What it is | Who uses it |
|---|---|---|
| Holarc | The lifetime health record. Every fact carries its source, its date and how far that source has earned trust. It reads almost any format: laboratory results, hospital records, PDFs, faxes, phone photographs of a report, imaging. | Everything else writes to it and reads from it. Patients never see it; their guide knows it. |
| My Talisman | The patient’s own app, with a guide to talk to or type to. It answers medical questions with referenced evidence, carries the whole family, and connects to any doctor’s portal. It informs; it never diagnoses or prescribes. | Patients, parents, caregivers |
| Arca | The clinic system you talk to: medical records and practice management run by a voice guide, with people approving the work. One system for every outpatient specialty. | Doctors, nurses, front desk, billers, practice owners |
| Keelson | The imaging platform: archive, viewer, reporting and AI support for X-ray, CT and MRI. | Imaging centers and radiology groups |
| NERD | Numerical Evidence and Research Discipline. The evidence engine: it scores every claim, runs the research, and returns a verdict a person can act on. | Researchers, clinics, and the guide inside every product |
| Finarc | The lifetime financial record and a financial adviser that does not sleep, for people, small businesses and trusts. Every account, asset, policy and obligation in one place, with walls between entities so each can be audited on its own. | Families, small business owners, and the banks, lenders and insurers that serve them |
| The foundation | One sign-in, one encryption system, one gateway to outside AI, one record engine, and tamper-evident logs under every product. | Everyone, invisibly |
One repository under everything
Every application in the suite reads from and writes to the same record. Arca does not keep a chart of its own, Keelson does not keep a separate patient index, My Talisman does not hold a copy of the history on the phone. Each produces facts and hands them to the record as events; each reads back the views it needs. There is nothing to reconcile between products because there is one log.
The same repository is what licensees buy. A laboratory network, a device maker, a specialty clinic chain, a lender or an insurer can run its records in a partition of Holarc or Finarc, with its own encryption keys, its own trust settings, its own additions to the data type registry and its own views. The licensee sees its own data and nothing else, and the walls are the ones the records already enforce between custody domains and financial entities. Inside those walls it gets the append-only log, the hash chain with outside anchors, the registry, the learned weights with their basis channels, lawful erasure by key destruction, the de-identified research path, and the post-quantum protection described in Paper 4. Where a person in a licensee’s partition also holds a personal record, the two link only through the identity vault and only when that person authorizes it. Paper 1 sets out the rules.
Why they belong together
Each piece is saleable alone and worth more beside the others, for a reason specific to this design.
Every piece produces outcomes, and outcomes are what the record learns from. A clinic system records what a treatment did. A patient app records how someone felt afterward. An imaging platform records what a reader found and what later proved true. A financial record records whether a statement reconciled. Those are the independent judgments Paper 2 shows the trust model needs and cannot generate for itself. Paper 2 also shows what happens to a record that tries: a weight carried by the system’s own agreement climbs toward certainty on no evidence at all, and has to be capped.
A record surrounded by applications has a supply of independent judgments. A record sold on its own does not, and its weights sit near their starting figures indefinitely. That is the argument for the suite. It is a claim about where hard labels come from, and bundling has nothing to do with it.
The second reason is narrower and holds regardless: one sign-in, one encryption system and one record engine built once instead of six times.
The one guide
One guide service runs both the patient’s guide and the clinic’s guide, with different tools, permissions and voices. Patients never see the names Holarc, NERD or Keelson. The evidence score on every answer My Talisman gives a patient, and on every evidence question Arca’s guide answers for a clinician, comes from NERD.
Voice is never mandatory. Every voice interaction can be replaced by typing. The guide is a feature of every product and carries no price or market of its own.
The line that does not move
Across the suite, AI drafts, retrieves, checks and informs, and licensed people diagnose, prescribe and sign. My Talisman never diagnoses, never prescribes beyond over-the-counter products, and never judges a photograph of a skin spot. NERD advises and people decide. The line is a design constraint, and several of the regulatory questions below exist because of where it sits.
The markets
Published estimates for 2025 and 2026. Research firms define these markets differently, so each is given as a range between the low and high estimate. Several of them overlap, and adding all of them together would count some spending twice.
| Segment | Products | Global, per year | United States, per year |
|---|---|---|---|
| Electronic health records | Arca, Holarc | $30.3B to $37.5B (2026) | $9.4B to $15.0B; ambulatory alone $4.05B (2026) |
| Practice management | Arca | $17.2B to $18.8B (2026) | $5.1B (2026) |
| Patient engagement and personal health records | My Talisman | $33.8B to $41.2B (2026) | $14.6B (2025) |
| Imaging archive and reporting | Keelson | about $6.8B to $7.1B (2025 to 2026) | not published at this scope |
| Real-world evidence for research | NERD, Holarc | $2.6B to $6.2B (2025 to 2026) | about $1.0B (2025) |
| Financial records: data aggregation, personal finance and small business accounting | Finarc | about $16.5B to $16.9B (2025) | not published at this scope |
| Healthcare analytics (adjacent) | NERD, Holarc | $78.3B to $81.9B (2026) | $28.2B (2026) |
| Hospital information systems (adjacent) | Arca Hospital | $63.8B (2024) | not separated |
| Open banking (adjacent) | Finarc | $29.8B to $35.7B (2025 to 2026) | $9.9B (2026) |
Adding only the core segments, which are records, practice management, patient engagement, imaging, research evidence and financial records, the suite addresses about $107 billion to $128 billion a year worldwide and about $30 billion to $36 billion a year in the United States. Healthcare analytics, hospital systems and open banking sit next to those numbers and overlap them in part.
The personal finance software line is small on its own, about $1.4B to $1.8B worldwide, because consumers have never paid much for a budgeting app. Finarc’s market is larger than that line, and the reason is in the business model below: it includes the financial businesses that pay to reach and serve verified customers, which is why open banking and data aggregation belong beside it.
These figures are drawn from a pooled set of published estimates and no single commissioned study, and we do not map each figure to a single firm because the underlying sources define their segments differently. The firms behind them are Fortune Business Insights, Grand View Research, Precedence Research, Mordor Intelligence, Towards Healthcare, MarketsandMarkets, Global Market Insights, Expert Market Research, Persistence Market Research, Nova One Advisor, IMARC and Market Research Future. A reader who needs one figure attributed should take it from the firm’s own publication.
Counting customers instead of dollars
Market sizes are estimates built on other estimates. Counts of institutions are harder to argue with.
| Who | How many | Source |
|---|---|---|
| US physician group practices | 395,000 | Definitive Healthcare |
| US imaging centers | 15,000 | Definitive Healthcare |
| US urgent care centers | 11,000 | Definitive Healthcare |
| US ambulatory surgery centers | 6,436 | MedPAC, March 2026 |
| US direct primary care practices | more than 2,700, and growing | DPC Frontier via Atlas.md |
| Hospitals in the Gulf states | about 882 | GCC Statistical Centre |
To show the scale with one assumption: if Arca averaged $6,000 a year per practice, which is about $500 a month and inside what small practices pay today, the 395,000 US group practices alone would be a $2.4 billion yearly market for Arca. That figure is an assumption and is identified as one everywhere it appears.
What each piece sells into
Arca, the clinic system
One code base serves every outpatient specialty through skins that set templates, order sets, billing rules and dashboard widgets. The build starts with cash-pay clinics, where the rules are lightest and the guide can prove itself fastest.
Two design decisions are commercial arguments as much as technical ones.
Physicians rate their electronic health records at 45 out of 100 on the standard usability scale, which is a failing grade and below Excel. Our target for every role is 90, and no release ships if any role tests below 80. Nobody has to sit through training: the guide can do or teach every function in the system, and a function does not ship unless a test proves the guide can teach it.
Drug interaction alerts are overridden about 90 percent of the time in published studies. In Arca each clinician chooses whether an alert interrupts, sits in a side panel, goes to a daily digest, or stays off. A short list of true safety alerts stays on because the clinic carries the liability. Any alert that doctors accept less than 30 percent of the time stops interrupting on its own.
My Talisman, the patient’s app
Every patient of an Arca clinic gets My Talisman, and sign-up needs no app: the patient opens an invitation link in any browser. The two are deliberately independent. My Talisman never depends on Arca, and Arca never requires a patient to install My Talisman.
The argument for it is the record’s argument turned around. Most systems treat the patient as the last person to ask. A lifetime record that reads any format, carries the whole family and connects to any portal makes the patient the most complete source in the system. It is also, in the language of Paper 2, a supply of hard labels: a patient reporting what happened after a treatment is an independent judgment, which is the scarce input the trust model needs.
Keelson, the imaging platform
An outpatient imaging center, a radiology reading group or a specialty clinic with its own scanners can run on Keelson as its full archive, viewer, reporting and workflow, under its own name, priced per study. The imaging center skin in Arca is Keelson’s radiology workflow, so scheduling, billing and the patient portal are built once.
The regulatory position is specific. The archive, sharing and workflow need no FDA clearance. The diagnostic viewer needs a 510(k) before primary reading, and so does each AI tool that detects, triages or measures.
NERD, the evidence engine
NERD scores a study from 0 to 100 on six dimensions and returns a verdict of one word: Pursue, Probe, Pivot or Park. Null results are flagged, kept, and earn the same credit as positive ones. Claims supported mostly by inherited evidence stay under a ceiling until they are tested directly.
Conventional care, alternative care and our own products are scored by the same rules. That is a commercial decision as much as an intellectual one. It keeps the record’s standing as an honest broker, and it protects our own products from the charge that their evidence was graded on a curve.
NERD sells research answers to universities, drug makers and device makers. It never sells the raw record. Researchers pay for answers, and the records stay where they are, which is what keeps patients, clinics and regulators on the same side of the question.
NERD is also where the population-level learning belongs. Paper 1 explains why both records leave the population level of the trust model switched off: with it on, a source’s weight can move because some other source was judged, and the record can no longer explain why. Learning across records, and across licensee partitions, is a consented operation, and the consent gate is NERD’s.
Finarc, the financial record
Finarc earns from two directions. The user pays for an adviser that knows everything about their money. The banks, lenders, insurers and advisers who serve that user pay to work from a complete, verified picture, with the user’s permission.
The two-sided shape is the lesson of the segment. Budgeting apps see a slice of a person’s money and people will not pay much for a slice; Intuit shut down Mint in 2024. Intuit paid $7.1 billion for Credit Karma in 2020 for one piece of the second side of this model, and Plaid was valued at $8 billion in 2026 for the bank connections alone.
Finarc is also where the trust model has its cleanest supply of independent labels, and Paper 2 explains why that matters. A statement that closes, a transcript that arrives from a tax authority, or an instrument that clears is ground truth that did not come from the system’s own answer. A health record waits years for an outcome; a financial record gets one every month.
Finarc’s tax filing, investment advice, audited figures and insurance features each launch with the license or licensed partner they require.
Licensed partitions of the record
Licensing is the seventh product in practice, though it sells the foundation itself. A company that licenses a partition brings its own data and keeps it to itself. It does not buy Arca or My Talisman, and its patients or customers never see our names unless it chooses to connect them to the apps.
What it gets is the part of a record system that is hardest to build and easiest to get wrong: permanent, hash-chained storage anchored outside the database; a data type registry that admits new kinds of data without a schema change; provenance and a learned weight on every value; lawful erasure that leaves the proof intact; identity kept apart from records; and encryption that keeps a stolen copy unreadable after large quantum computers arrive. The licensee’s trust settings, keys, registry additions and views are its own. Its weights learn from its own outcomes, and nothing it stores is visible to any other tenant or to us beyond what operating the service requires and the license permits.
The natural licensees are the companies that already hold longitudinal data and have no good place to keep it: laboratory networks, device and diagnostic makers, specialty clinic chains, imaging groups that do not want Keelson under their own name, lenders and insurers that need a verified position over time. For a device maker in particular, a partition turns every treated patient into a long-term outcome record, scored by rules the device maker does not control, without a record system or a compliance program of its own.
Attached technologies and the device outcome loop
A device or diagnostic company can connect to the record under six commitments: identify the patient at the point of care, send every result in, read back what it needs, keep the method, be scored like everyone else, and reach people through the apps.
The fourth commitment is what makes the arrangement acceptable to a device maker. Algorithms, settings logic and reconstruction methods stay on the company’s own servers. The record stores the result of using a method, never the method.
| Technology | Sends in | Reads back | Keeps |
|---|---|---|---|
| Navikara (Quantumis Bio) | Intake findings, injury and inflammation data, treatment sessions, required photographs and gait video, reported response | History, medications, prior imaging | Settings logic and each patient’s treatment curve |
| V-HDI CT and Concordance Volumetric Tomography (Quantumis Bio) | Reports; the studies go to Keelson in DICOM | Prior studies for comparison | Reconstruction and processing methods |
| Quantumis Bio DICOM viewer | Measurements, segmentations, reprocessed series | Studies from Keelson | The viewer software |
| CardioScore MCG | Test results, index scores, reports | Cardiac history, medications, prior tests | Scoring algorithms and its historical database |
| MicroDose | Treatment sessions, glucose and metabolic labs, outcomes | Conditions, medications, labs | Treatment protocols |
For a device company the connection turns every treated patient into a long-term outcome record without building a record system or a compliance program, and produces evidence scored by rules the company does not control. The fifth commitment protects both sides.
Quantumis Bio is a separate company of which Greg Stoutenburgh is chief executive, and Navikara is its product. The commercial terms between the two companies are at arm’s length and are documented separately.
How it earns
Six streams. Every figure attached to them is an assumption, identified as one, and none of them is a quotation or a signed price.
| Stream | Unit | Assumed figure |
|---|---|---|
| Clinic subscriptions for Arca and Keelson | per practice, per center | $6,000 a year per practice; $24,000 a year per imaging center |
| User fees for Finarc and premium features in My Talisman | per user | $10 a month for Finarc |
| Fees from financial businesses to serve Finarc users from verified data, with each user’s permission | per user, per year | about $40 a year per Finarc user |
| Research answers from NERD | per engagement | varies by scale; never the raw record |
| Licensed partitions of the record for other companies, each seeing only its own data | per licensee, by volume | varies by scale |
| Device access | per connected device maker | no figure set |
Two rules bound all of it. Records are never sold, rented or transferred, and partners pay for the platform and never per record or per person. We never sell patient data. A licensee’s data is the licensee’s, and it leaves the partition only by that licensee’s instruction and the consent of the people in it.
Competition
Epic holds most large US health systems and about one in five outpatient sites. It sits at the top of user satisfaction ratings among the systems that exist, and it is also the target of the most complaints: clicks, note bloat, an inbox that follows doctors home, and cost. Texas sued Epic in December 2025 alleging it blocks competitors and overcharges. Epic and athenahealth now give away AI-drafted visit notes as part of their systems, and Oracle markets its new record as voice-first.
The incumbents are not standing still, and the obvious features are being commoditized as we build. An AI scribe is no moat. What the incumbents are not doing, as far as any public material shows, is weighing the evidence behind each fact, learning which sources deserve trust, recording where each judgment came from, and declaring which sources share an upstream. That is the claim the suite rests on, and Paper 2 is where it is argued and measured.
The size of the claim is measured there too, under a disclosed cold-start protocol. Against a source carrying a hidden bias the record holds its stated confidence in every run at every size while all three conventional methods’ confidence collapses to zero, identifies which source it is in 99 percent of runs by 250 records and every run from 500 on, and cuts the error of the resulting estimate by a factor of 96, landing within half a percent of the best weighting of the surviving sources. Against a source that has quietly broken it cuts wrong facts by 64 percent, which is 99.7 percent of everything available to be won. Against ordinary independent mistakes it gains nothing, and neither does weighting every source by its true reliability, so there was nothing there to gain. Those are failure modes every large record system has and none of them can see. A competitor could copy the scribe in a quarter. Copying a record that knows which of its feeds broke eighteen months ago requires having kept the evidence, which is a decision made years earlier.
Who buys first, and in what order
We enter where the incumbents are weakest and the rules are lightest: cash-pay practices, direct primary care, concierge and wellness clinics. They do not bill Medicare, so they need no federal certification to buy, and the systems they use today offer thin AI at $290 to $770 a month. From there the 395,000 group practices, the 15,000 imaging centers, the 11,000 urgent care centers and the 6,436 surgery centers follow as federal certification and payer connections come online. Hospitals are not the first target and the system is built to run them. Licensed partitions run in parallel with the clinic entry, because the first licensees are companies that already hold data and need a place to keep it.
Outside the United States, in order: English-speaking markets with private, cash-pay clinics, which are the United Kingdom’s private sector, Australia and Canada; private hospital groups in the Gulf states, India and Latin America, where a full incumbent installation is out of reach on price; the European Union, once its new certification for health record systems takes effect from March 2027; and national public-sector systems last.
What stands between here and revenue
Every market figure above depends on this list being short.
| Before this can be sold | Status |
|---|---|
| A real deployment with real records | None. The suite runs on invented data |
| The first calibration of learned weights against real outcomes | Follows the first deployment, and will be published before sale |
| Federal EHR certification, for any clinic that bills Medicare | Not begun. Not required for the cash-pay entry market |
| A 510(k) for the diagnostic viewer before primary reading, and for each measuring AI tool | Not begun |
| IRS e-file authorization, investment adviser registration, CPA attestation, insurance licensing for Finarc’s regulated features | Each gated behind its own license or licensed partner |
| An operating software team | Not yet hired |
| Per-person encryption keys in dedicated hardware | Designed, not yet in place |
| Commercial license terms and tenant onboarding for licensed partitions | Drafting |
This paper carries no valuation, no forecast and no revenue projection. The valuation work, with its methods and its assumptions, is a separate document with a separate purpose.
Sources
Market sizes are drawn from published estimates by Fortune Business Insights, Grand View Research, Precedence Research, Mordor Intelligence, Towards Healthcare, MarketsandMarkets, Global Market Insights, Expert Market Research, Persistence Market Research, Nova One Advisor, IMARC and Market Research Future. Institution counts are from Definitive Healthcare, MedPAC (March 2026), the American Hospital Association, DPC Frontier via Atlas.md, the British Medical Association and the GCC Statistical Centre.
- Physician EHR usability on the System Usability Scale: Melnick et al., Mayo Clinic Proceedings, 2019
- Software usability benchmarks: MeasuringU
- Drug interaction alert override rates: JAMIA, 2012
- Texas v. Epic Systems, filed December 2025
- Intuit’s acquisition of Credit Karma (2020) and the closure of Mint (2024), from company announcements
- European Health Data Space regulation, certification for health record systems from March 2027
- Stoutenburgh G. The Evidence-Weighted Record. CS Quantum Suite White Paper Series, Paper 1, Rev 6, October 7, 2026
- Stoutenburgh G. The Mathematics. CS Quantum Suite White Paper Series, Paper 2, Rev 6, October 7, 2026
- Stoutenburgh G. Data Security in the Quantum Era. CS Quantum Suite White Paper Series, Paper 4, Rev 5, October 7, 2026
© 2026 VitaNexus Holdings, Inc. All rights reserved. Confidential. Holarc, Finarc, Arca, Keelson, My Talisman and NERD are working names pending trademark clearance. Every pricing figure in this paper is an assumption and is identified as one. The suite has no customers, no revenue, and holds no real patient or financial record.
The white papers · Paper 4 of 4 · Rev 5 · October 7, 2026
Data Security in the Quantum Era
How the Record Stays Private for a Lifetime.
The full paper, exactly as published. By Greg Stoutenburgh; founders and inventors Greg Stoutenburgh and Faiz Chowdhury. Also available as PDF and Word on request.
Summary
Health data is stolen more than almost any other kind of data, and it stays sensitive for the rest of a person’s life. Financial data is stolen for immediate gain and exposes everything a person owns and owes. A lifetime record of either kind concentrates exactly what thieves want, and this suite keeps both.
The design goal can be stated precisely: a stolen copy of either record is ciphertext, and it stays ciphertext after large quantum computers arrive. A stolen key opens one person’s record, never the database.
Five design choices deliver that goal.
- Nothing stored depends on public-key cryptography. Stored data is encrypted with a 256-bit symmetric cipher in an authenticated mode, and the keys that protect it are wrapped with other symmetric keys inside dedicated hardware. RSA and elliptic-curve algorithms, the ones a quantum computer breaks, appear nowhere in the storage path. NIST expects AES-256 to remain safe for a very long time, quantum computers included.
- Keys are separated by subject. A stolen or cracked key exposes one record, or one entity’s books, or one licensee’s partition. Destroying a key erases that subject’s data everywhere, backups included, by a method NIST recognizes as a purge.
- Identity, records, clinic finances and research live apart. Each sits in its own store under its own keys, so no single theft yields a named, complete record of either kind.
- Connections use post-quantum key exchange. Traffic recorded today cannot be decrypted later, because the key exchange combines a classical algorithm with ML-KEM, the NIST post-quantum standard.
- The records verify their own integrity anchors. Each record checks the Ed25519 and ML-DSA-87 signatures on its outside anchors itself. It does not take the anchoring service’s word that a signature is valid, because a record that trusts the service that signs it has no defense against that service.
Under HIPAA, health information encrypted to NIST’s standard with uncompromised keys is not “unsecured,” because HHS guidance treats it as “unusable, unreadable, or indecipherable to unauthorized persons.” Under the FTC Safeguards Rule, which governs a non-bank financial company such as Finarc, encrypted customer information counts as unencrypted only if its key was taken too. Both records are designed to those standards from the first record.
We also state what the design cannot do. An attacker who takes over the running application can read what that application can read while the compromise lasts. The design limits how much that is, records every access, and never claims that de-identified data is anonymous.
The papers in this series
- Paper 1, The Evidence-Weighted Record. What the record is and how it works.
- Paper 2, The Mathematics. The proofs and the measurements.
- Paper 3, Markets and Applications. The products, the markets, and the gap to revenue.
- Paper 4, this document. The security design and its limits.
The exposure
In calendar year 2024, HHS received reports of 663 breaches of unsecured protected health information that each affected 500 or more people, with 242.9 million individuals affected across those reports (HHS Office for Civil Rights). The largest single event, the ransomware attack on Change Healthcare, affected 192.7 million people by the company’s final count.
A stolen credit card can be canceled. A diagnosis, a genome, a psychiatric history or a record of pregnancy cannot. Much of a health record stays sensitive for the patient’s lifetime, and genetic information also bears on parents, siblings and children who never consented to anything.
A lifetime financial record carries a different shape of the same problem. Account numbers can be changed and balances move, so the half-life of any single datum is shorter. What does not change is the map: who owns what, through which entities, with what obligations, and where the money moves each month. That map is what enables targeted fraud, and it is exactly what a record built to tie a family’s entities together holds.
Conventional databases of both kinds were not built for this. Most protect the whole database with a small number of keys or credentials held by the application and its administrators. When one of them is stolen, every record behind it is exposed at once.
The quantum problem
Almost all data protected in transit today, and much of the data protected at rest, depends on public-key algorithms: RSA and elliptic-curve cryptography. Their security rests on mathematical problems that a large quantum computer running Shor’s algorithm solves efficiently.
How close that is. In 2019 the best published estimate for breaking 2048-bit RSA was about 20 million noisy qubits running for 8 hours. In May 2025, Craig Gidney of Google Quantum AI published an estimate of fewer than one million noisy qubits in less than a week (Gidney, 2025). The Global Risk Institute’s 2025 survey of quantum experts found that a quantum computer able to break RSA-2048 within 24 hours is “quite possible” within 10 years (28 to 49 percent) and “likely” within 15 years (51 to 70 percent) (Mosca and Piani, 2026).
Harvest now, decrypt later. An attacker does not need a quantum computer today to exploit one tomorrow. Encrypted data can be copied now and decrypted when the machine exists. The Office of Management and Budget put it directly in 2022: “encrypted data can be recorded now and later decrypted by operators of a future CRQC” (OMB M-23-02).
Why a lifetime record is first in line. Michele Mosca framed the test as an inequality (Mosca, 2018). Let x be how long data must stay confidential, y how long it takes to migrate to quantum-resistant encryption, and z how long until a quantum computer can break today’s public-key algorithms. If
x + y > z,
data protected today is exposed during its useful life. For a lifetime record, x alone is measured in decades, longer than the expert estimates of z. Any such record sent or stored under RSA or elliptic-curve protection today should be assumed readable within the subject’s lifetime.
This failure is silent, in the same way as the binary cliff in Paper 2. Data protected by today’s public-key encryption looks safe today and becomes readable later, and nothing in the database signals the change.
The standards now in force
- Post-quantum standards. NIST published FIPS 203 (ML-KEM, key establishment), FIPS 204 (ML-DSA, digital signatures) and FIPS 205 (SLH-DSA, hash-based signatures) on August 13, 2024. NIST selected HQC as a backup key-establishment algorithm in March 2025, with a final standard expected in 2027. FIPS 206, based on Falcon, is still in preparation.
- Retirement dates. NIST IR 8547 (November 2024, draft) proposes that RSA and elliptic-curve algorithms at the 112-bit security level be deprecated after 2030 and disallowed after 2035.
- Federal agencies must migrate. An executive order of June 22, 2026, and OMB memorandum M-26-15 of June 24, 2026, direct civilian federal agencies to move priority systems to post-quantum key establishment by the end of 2030 and to complete migration by 2035.
- National security systems. The NSA’s CNSA 2.0 specifies AES-256, SHA-384 or SHA-512, ML-KEM-1024 and ML-DSA-87, requires web, server and cloud services to use them exclusively by 2033, and requires all national security systems to be quantum-resistant by 2035.
- Symmetric encryption stays. NIST’s position is that quantum computers are likely to give little or no advantage against AES, and that AES-192 and AES-256 will still be safe for a very long time.
For a platform that expects to work with federal agencies, federal health programs and defense health systems, these dates are procurement requirements in waiting. Both records are built to them now, before any real data enters, which costs far less than migrating a live system later.
The regulatory benefit
HIPAA. The breach notification rule applies to “unsecured” protected health information: information not rendered “unusable, unreadable, or indecipherable to unauthorized persons” through a method specified by the Secretary of HHS (45 CFR 164.402). HHS guidance names two such methods (74 FR 42740): encryption consistent with NIST SP 800-111 for stored data and validated protocols in transit, where the key has not been breached, with decryption keys stored separately from the data they protect; and destruction consistent with NIST SP 800-88. When data meeting this standard is stolen and its keys are not, the theft is not a breach of unsecured protected health information.
HHS has proposed updating the HIPAA Security Rule to require encryption of all electronic protected health information at rest and in transit, with documented exceptions (90 FR 898, January 6, 2025). The rule is not final; the most recent regulatory agenda lists final action for July 2027. The design meets the proposed requirement today.
The FTC Safeguards Rule. 16 CFR Part 314 applies to non-bank financial institutions, which include tax preparers, lenders and companies that handle consumers’ financial data on their behalf. It requires encryption of customer information at rest and in transit unless encryption is infeasible and a qualified individual approves alternative controls. Since May 2024 it has required notice to the FTC, no later than 30 days after discovery, of any event involving unauthorized acquisition of unencrypted customer information of at least 500 consumers. Encrypted information counts as unencrypted if its key was also accessed.
State law follows the same pattern. California’s statute applies to unencrypted personal information, and to encrypted information when the key or credentials to decrypt it were also acquired.
The common thread across all three is that the key must not travel with the data. That is a storage architecture decision, and it is the one the next section describes.
How stored data is protected
One design rule: no public-key cryptography in the storage path
Every stored byte (events, documents, imaging studies, backups) is encrypted with a 256-bit symmetric cipher in an authenticated mode, which detects tampering with the ciphertext as well as hiding its content. The keys are arranged in a hierarchy:
- Data keys. Data is encrypted under keys scoped to the subject and the purpose: in Holarc, one per custody domain, which is the clinic’s chart and the patient’s personal record; in Finarc, one per entity, so the audit walls between a household, a business and a trust are walls in the cryptography as well as in the access rules. A licensee’s partition carries its own data keys on the same principle, so one tenant’s data is never readable under another’s key.
- Key-encryption keys. Data keys are stored only in wrapped form, encrypted under key-encryption keys with AES key wrapping.
- Hardware roots. Key-encryption keys never leave dedicated hardware security modules validated to FIPS 140-3, or an equivalent cloud key service, on systems separate from the data stores.
At no point does the storage path use RSA or elliptic-curve cryptography. A thief who copies every disk, database and backup and waits for a quantum computer gets ciphertext protected only by AES-256, which a quantum computer does not meaningfully weaken.
One implementation of the ciphertext format serves both records, in the shared engine. One implementation means one place to get an authenticated mode right and one place to audit.
Keys separated by subject
A conventional database is protected by a few keys. These records are protected by one set per subject. A key that leaks, through a software flaw, a misconfigured backup or an insider, exposes one subject’s data. An attacker who wants the whole database has to break the hardware that guards the key-encryption keys.
Erasure by destroying a key
Destroying a subject’s data key makes every copy of that data unreadable at once, including copies in backups and archives nobody can reach to delete. NIST recognizes this method, cryptographic erase, as a purge-level sanitization technique (NIST SP 800-88 Rev. 2, section 3.2, 2025).
Both records use it to honor an erasure request without rewriting a permanent event log, and both record on the chain that the erasure happened. Charts a clinic is legally required to retain sit in the clinic’s custody domain under the clinic’s keys and are unaffected.
One consequence is specific to an evidence-weighted record. The judgments that moved a source’s reliability are kept as rows carrying the weight before and after, so erasing a person’s record leaves the learned reliability of its sources intact and auditable. Erasing a subject does not erase the fact that a feed was found unreliable.
Separation of what a thief would need to combine
A record is most dangerous when a name is attached to it.
- Identity vault. Names, birth dates and outside identifiers live in a separate store under separate keys. Records carry only random tokens.
- Clinic ledgers. Each clinic’s charges, claims and payments live in a separate database under that clinic’s keys, with no path to research, to NERD or to any other clinic.
- Research store. De-identified research records carry a different token again. Outside researchers never receive records; analyses run inside the walls and results are checked before release.
- Licensee partitions. A licensee’s data sits behind its own keys and its own access grants. No path exists from one partition to another, to the suite’s own records, or to research, except through the identity vault and the consent engine when a person authorizes a link.
- No payment authority. Finarc reads, drafts and instructs. A thief who controls Finarc cannot move money through it.
Data in motion, and the record’s integrity
Post-quantum connections from the start
Data in transit is the main target of harvest-now, decrypt-later collection, because an attacker can record it without breaking into anything. Connections use TLS 1.3 with hybrid key exchange, combining a classical elliptic-curve exchange with ML-KEM (FIPS 203). The session key stays secret if either algorithm holds, so the connection is no weaker than today’s standard and is protected against a future quantum computer. Links between servers and backup sites use ML-KEM-1024, the level CNSA 2.0 specifies.
An integrity chain that quantum computers do not break
Each event carries a cryptographic fingerprint of itself and of the event before it, so altering any past event breaks every fingerprint after it. The chain uses SHA-384. Quantum computers offer at most a square-root speedup against hash functions, which this output size absorbs, and the NSA specifies the same functions for national security systems through the quantum transition.
Every hour the chain head is anchored outside the database in write-once storage, with the chain’s Merkle root and size, and signed with both Ed25519 and ML-DSA-87 (FIPS 204). Several RFC 3161 timestamp authorities are queried in parallel so the anchor carries independent attestations of its time. The Merkle tree follows RFC 6962, so a single record can be proved to be in the chain without disclosing the rest of it.
Verifying the anchor instead of trusting it
The purpose of anchoring outside the database is to defend against an attacker who has the database. A record that accepts the anchoring service’s word that a signature is valid has no defense against that service, or against anyone who has compromised it. So each record verifies the Ed25519 and ML-DSA-87 signatures on its anchors itself, under the anchor context string, with the engine’s own verification code.
The engine’s conformance suite includes two cases that exist only to prove this holds: a consistent truncation of the chain and a consistent rewrite of the last event. Both leave every link and the stored head in agreement, so the database alone cannot catch either. The suite requires that the local walk pass and that the anchor check fail, so a clean local verification is never read as a sound record.
Algorithms that can be replaced
Every key and every ciphertext is labeled with the algorithm and version that produced it. Because data keys are wrapped by key-encryption keys, replacing a key-encryption algorithm means rewrapping the small data keys without re-encrypting the stored data. If NIST revises a standard or a weakness is found, migration happens in the background without taking the record offline. Cryptographers call this crypto-agility, and the federal transition plans require it.
The same property applies to the chain. The chain encoding is versioned, and a version boundary lets the hashing advance without rewriting a stored fingerprint: older entries verify under the encoder that wrote them, and new entries chain onto them under the current one.
What a thief gets
| What is stolen | What the thief holds | Usable |
|---|---|---|
| Disks, database dumps or backups | AES-256 ciphertext; keys are elsewhere, in hardware | No, today or after large quantum computers arrive |
| Network traffic recorded in transit | Sessions protected by hybrid ML-KEM key exchange | No, today or later |
| One subject’s data key | That one subject’s record in one custody domain or one entity | One record, not the database |
| One licensee’s keys | That partition’s data | One partition; nothing from any other tenant |
| The identity vault | Encrypted names and identifiers with no health or financial data | No clinical or financial information |
| The research store | Encrypted, de-identified records under research tokens | Encrypted; treated as sensitive because re-identification is possible |
| Control of the running application, with live credentials | What that application can read during the compromise | Partly; limited by per-subject keys, access controls, logging, export approvals, and no payment authority |
| An insider’s legitimate access | What that role is permitted to see | Every read is logged on the chain and visible to the patient, owner or trustee |
What we do not claim
A security claim is only as good as its stated limits.
- A compromised running system is different from a stolen copy. Software that is running must be able to decrypt the records it serves. An attacker who controls that software can read what it reads while the control lasts. The design limits the damage with per-subject keys, role and purpose checks on every read, rate limits, alerts on unusual access, hardware-key sign-in for staff, and approval by named people for any large export. It does not claim such data is unusable.
- De-identified is not anonymous. Latanya Sweeney showed that 87 percent of Americans were unique on ZIP code, sex and birth date alone (Sweeney, 2000). Rocher and colleagues estimated that 99.98 percent of Americans could be re-identified from 15 demographic attributes (Rocher et al., 2019). A lifetime record is richer than either. The research store is encrypted, answers are released instead of records, and every result is checked so no small group can be singled out.
- New algorithms are younger. ML-KEM and ML-DSA have been studied intensively but for less time than RSA. Hybrid key exchange, and signing anchors with both a classical and a post-quantum algorithm, protect the design if a weakness in the newer algorithms is found.
- Endpoints belong to their owners. A patient’s phone or a clinic’s workstation can be compromised outside the record’s control. The design limits what any one device can reach and keeps device caches encrypted.
- Sign-in keys will migrate too. Today’s hardware security keys authenticate with elliptic-curve signatures. Sign-in happens in real time, so recorded traffic cannot be replayed later to break it, and post-quantum authenticators will be adopted as the standards for them are completed.
- The standards are still moving. HQC and FIPS 206 are not final, and NIST IR 8547 is a draft. Crypto-agility exists so the design can follow the final versions.
- A hash chain proves tampering, not correctness. It shows that what was written has not been altered. It says nothing about whether what was written was true when it was written, which is the subject of Papers 1 and 2.
Where the design stands
Everything described above is in place in both records, with two exceptions that bear on the claims. Data keys today are scoped per tenant and per purpose; the per-person and per-entity keys that make a stolen key open one record and never a tenant’s records are designed and are the next item to complete. Post-quantum signing of the anchors runs in software today; moving it inside the hardware security module follows. Until the first is complete, the statement that a stolen key opens one person’s record describes the design, and a compromised tenant key would expose that tenant’s records.
Both items, together with a written HIPAA risk analysis with named Privacy and Security Officers and an outside penetration test, come before any real record enters either system. HITRUST, SOC 2 and ISO 27001 certification follow the first deployment. Neither record holds a real patient record or a real financial account today.
Conclusion
Conventional records protect millions of subjects with a few keys and rely on public-key algorithms that quantum computers will break, so data stolen from them today should be assumed readable within their subjects’ lifetimes. These records keep public-key cryptography out of the storage path, separate keys by subject and by tenant, keep identity, records, finances and research apart, use post-quantum key exchange for every connection, and verify their own anchors instead of trusting the service that signs them.
A stolen copy is ciphertext, and it stays ciphertext after large quantum computers arrive. With per-subject keys in place, a stolen key opens one person’s record and never the database.
References
- Gidney C. How to factor 2048 bit RSA integers with less than a million noisy qubits. arXiv:2505.15917. May 2025.
- Gidney C, Ekerå M. How to factor 2048 bit RSA integers in 8 hours using 20 million noisy qubits. Quantum. 2021;5:433.
- Mosca M. Cybersecurity in an era with quantum computers: will we be ready? IEEE Security & Privacy. 2018;16(5):38–41.
- Mosca M, Piani M. Quantum Threat Timeline Report 2025. Global Risk Institute and evolutionQ; March 2026.
- NIST. FIPS 203, Module-Lattice-Based Key-Encapsulation Mechanism Standard; FIPS 204, Module-Lattice-Based Digital Signature Standard; FIPS 205, Stateless Hash-Based Digital Signature Standard. August 13, 2024.
- NIST. NIST IR 8545, and the selection of HQC. March 2025.
- NIST. NIST IR 8547 (Initial Public Draft), Transition to Post-Quantum Cryptography Standards. November 12, 2024.
- NIST. Special Publication 800-88 Rev. 2, Guidelines for Media Sanitization. 2025.
- NIST. Special Publication 800-111, Guide to Storage Encryption Technologies for End User Devices.
- NSA. Commercial National Security Algorithm Suite 2.0.
- Office of Management and Budget. M-23-02 (2022) and M-26-15 (June 24, 2026).
- HHS Office for Civil Rights. Breach Portal and annual reports to Congress.
- HHS. Guidance Specifying the Technologies and Methodologies That Render Protected Health Information Unusable, Unreadable, or Indecipherable to Unauthorized Individuals. 74 FR 42740 (2009).
- HHS. HIPAA Security Rule Notice of Proposed Rulemaking. 90 FR 898, January 6, 2025.
- Federal Trade Commission. Standards for Safeguarding Customer Information, 16 CFR Part 314, including the breach notification amendment effective May 2024.
- Laurie B, Langley A, Kasper E. Certificate Transparency. RFC 6962, 2013.
- Adams C, Cain P, Pinkas D, Zuccherato R. Internet X.509 Public Key Infrastructure Time-Stamp Protocol. RFC 3161, 2001.
- Sweeney L. Simple demographics often identify people uniquely. Carnegie Mellon University, 2000.
- Rocher L, Hendrickx JM, de Montjoye Y-A. Estimating the success of re-identifications in incomplete datasets using generative models. Nature Communications. 2019;10:3069.
- Crosby SA, Wallach DS. Efficient data structures for tamper-evident logging. 18th USENIX Security Symposium, 2009.
- Stoutenburgh G. The Evidence-Weighted Record. CS Quantum Suite White Paper Series, Paper 1, Rev 6. October 7, 2026.
- Stoutenburgh G. The Mathematics. CS Quantum Suite White Paper Series, Paper 2, Rev 6. October 7, 2026.
- Stoutenburgh G. Markets and Applications. CS Quantum Suite White Paper Series, Paper 3, Rev 6. October 7, 2026.
This paper supersedes Holarc Data Security in the Quantum Era and Finarc Data Security in the Quantum Era, and covers both records on the shared engine.
© 2026 VitaNexus Holdings, Inc. All rights reserved. Holarc, Finarc and NERD, including their software, designs and documentation, are the property of VitaNexus Holdings, Inc. Confidential.
THE COMPANY
Founders and inventors
Greg Stoutenburgh
Founder and Inventor
Greg founded a medical imaging company and invented its Volumetric High-Definition Imaging technology, conceiving the system from the first sketch and designing the defining characteristics of the major iterations of both its platforms over fourteen years with the company, five of them as chief executive. Its devices for human use carry FDA clearances and CE marks. He is lead inventor on 24 issued patents and oversaw a company portfolio of more than 106 patents issued worldwide, and he has led or taken part in more than twelve mergers and acquisitions.
He was an operator first. He founded two management services organizations for physician and dental clinics and several veterinary clinics, and has run, managed or consulted for well over a hundred medical, dental and veterinary practices, which are the practices this suite is built to serve. He is Chief Innovation Officer of Graphene Valley Corporation and leader of the GVC Innovations Lab, and chief executive of Quantumis Bio, which develops regenerative and diagnostic products for human medicine.
His coursework and degrees are in biology, chaos and complexity, chemistry and philosophy. He is the author of Actual Intelligence, The Subject Answers Back and Inside the Walls, and of two children's books, The World Answers Back and Keep an Eye on Grandpa. He founded the Human Continuity Project and has served on the board of trustees of Casa Romantica since 2013. He conceived and created NERD.
Faiz Chowdhury
Founder and Inventor
Faiz is the founder and largest shareholder of Graphene Valley Corporation and owns and runs several other United States companies. He is co-inventor of the inventions in the Quantavera suite.
How the suite was built, said plainly
The founders conceived the inventions, set every design requirement and made every decision. They used AI to do the work that teams of programmers and mathematicians would otherwise do, under their direction, because it was the more efficient way to execute. The method is Actual Intelligence, the arrangement of people and AI that Greg Stoutenburgh defines in his book of the same name, and the discipline that makes it work is encoded in NERD. That is a permanent cost advantage on every product that follows, and it is described in full under Actual Intelligence.
Contact: Greg Stoutenburgh · 714-402-0455 · greg@graphenevalley.com
THE COMPANY
Ownership, names and copyright
Clean title, documented conception, and a transfer that completes before the first draw.
Who owns the suite today
The software, designs and documentation of the Quantavera suite are the property of VitaNexus Holdings, Inc., a Wyoming corporation wholly owned by Greg Stoutenburgh. Quantavera, Inc. is being formed to hold the suite, and the assignment transfers it on formation, before the first draw of this round. Nothing in the suite is owned by, licensed from, or encumbered by any other company either founder is involved with.
The conception record
A dated conception and authorship record documents each invention: what it is, who conceived it, when, and what evidence establishes the date. It is maintained to the standard a patent office asks for, and it is available in diligence. Current Patent Office guidance treats AI as a tool and looks to human conception, which is the standard the record is written to.
Names
Holarc, Finarc, Arca, Arca Vet, Keelson, My Talisman, Pet Talisman, My Helm and NERD are working names pending trademark clearance, and this round funds clearance and filing across the product line with international coverage. The NERD mark keeps its own established glasses device. Product symbols are used with their wordmarks as lockups rather than alone.
Copyright
© 2026 VitaNexus Holdings, Inc. All rights reserved.
REFERENCE
Sources and notices
Market sizing
Segment figures are pooled from published 2025 and 2026 estimates by Fortune Business Insights, Grand View Research, Precedence Research, Mordor Intelligence, Towards Healthcare, MarketsandMarkets, Global Market Insights, Expert Market Research, Persistence Market Research, Nova One Advisor, IMARC and Market Research Future. Research firms define these segments differently, so each figure is given as a range. Anyone needing a figure attributed should take it from the publishing firm's own report.
Institution counts
United States physician group practices, imaging centers and urgent care centers: Definitive Healthcare. Ambulatory surgery centers: MedPAC, March 2026. Direct primary care practices: DPC Frontier via Atlas.md. Veterinary establishments and veterinarians: the American Veterinary Medical Association and the United States Census. Veterinary businesses and sector revenue: IBISWorld. Veterinary care and product spending: the American Pet Products Association. Pet insurance participation: the North American Pet Health Insurance Association. Hospitals in the Gulf states: the GCC Statistical Centre.
Valuation and cost references
Series A valuation medians: PitchBook-NVCA Venture Monitor, first quarter 2026, as reported by WaveUp, Series A Benchmarks 2026. Electronic health record development cost: Groovy Web, EHR Software Development Cost in 2026. Replacement cost is the company's own estimate on the stated assumptions, pending an independent appraisal.
Security and cryptography
Breach counts: United States Department of Health and Human Services, 2024. Quantum resource estimates: Craig Gidney, Google Quantum AI, May 2025, against the 2019 estimate it revises. Re-identification: Latanya Sweeney. Post-quantum standards: the National Institute of Standards and Technology.
Mathematics
The optimality bound for linear weighting: A. C. Aitken, 1935. The measured comparisons, the simulation protocol and the figures behind them are published in full in the company's white paper series, which states the protocol, the sample sizes and the spread for every figure quoted in this material.
Notices
Every figure attached to a revenue stream or a financial projection is an assumption, identified as one, and no figure in this material is a quotation, a signed price, a forecast or a promise of any result. The company has no customers and no revenue. No real patient record, clinic, client or financial account has entered any part of the suite; everything described runs on invented data with outside networks simulated. This material is for discussion with qualified prospective investors and is not an offer to sell or a solicitation of an offer to buy securities.
© 2026 VitaNexus Holdings, Inc. Product names are working names pending trademark clearance.