When Health Data Becomes a Decision: Wearables, AI, and Human Judgment

By Mustafa Kemal Calik, MD

Published 01-09-2026

A blood-pressure cuff, smartwatch, phone alert, ECG strip, and handwritten notes arranged around a decision point. When Health Data Becomes a Decision: Wearables, AI, and Human Judgment
Conceptual editorial illustration; not a real clinical record or patient encounter.

The patient places his telephone on my desk before I ask for it.

He has come prepared.

His blood-pressure readings are there, along with his resting heart rate, step count, sleep score, several single-lead ECG recordings from his smartwatch, and a graph showing that one measure of “recovery” has drifted downward during the past two weeks.

Then he opens another screen.

An AI tool has summarized the pattern for him.

Nothing on the screen looks absurd. Some of the information is useful. A few findings deserve another look. Others may be accurate without being clinically important. One measure comes from a device I did not prescribe and whose algorithm neither of us can meaningfully inspect.

He has not come because he is panicking.

He has come because he is trying to be responsible.

He looks at the graphs, then at me.

“Which of these should change what we do?”

That is the question.

Technology can measure more than ever before. Artificial intelligence can increasingly recognize patterns inside those measurements.

The harder problem begins afterward.

A number, signal, prediction, or recommendation has entered the room.

Now we have to decide how much weight it deserves.

The screen can be precise while the decision remains uncertain

Modern health technology can create an extraordinary feeling of precision.

A heart rate arrives as a number. Blood pressure appears beside a trend line. Sleep is divided into stages. An ECG trace crosses the screen with a label beneath it.

Precision is reassuring.

But precision of display can make one distinction surprisingly easy to forget:

A precise output is not necessarily a precise answer to the clinical question in front of us.

The U.S. Government Accountability Office’s 2026 assessment of wearables in clinical decision-making describes genuine opportunities. Wearable devices can collect information beyond clinical encounters, help establish longitudinal patterns, support earlier detection, and potentially contribute to more personalized care.1 The same assessment identifies important limitations. Devices vary in reliability, data can be difficult to integrate into clinical work, and the place of wearable information within medical decisions remains unevenly defined.

None of that is an argument against wearables.

A blood-pressure record at home may reveal a pattern hidden by one office measurement. Continuous glucose monitoring can make physiology visible between laboratory tests. A wearable ECG may capture an arrhythmia that has disappeared by the time a patient reaches the clinic.

These are real gains.

But information is only the beginning.

Accuracy is not the final test of a health tool. The harder test is whether it improves the decision that follows.

There are things machines can see that we cannot

I do not want to begin with the warnings.

There is a reason medicine wants these tools.

Human attention is powerful.

It is also finite.

Medicine increasingly generates more images, waveforms, laboratory values, records, messages, and published evidence than any individual clinician can continuously absorb. A machine can search patterns across quantities of information that would exhaust human attention long before the task was finished. A sensor can observe for days where a consultation sees minutes.

That difference matters.

AI is already being used in cardiovascular medicine for image and ECG interpretation, risk prediction, workflow support, case identification, and clinical decision support. A 2026 systematic review identified 32 randomized controlled trials of contemporary machine-learning and deep-learning interventions in cardiology, with 27 included in meta-analysis. Across heterogeneous applications, the authors found gains in workflow efficiency, patient engagement, and several clinical outcomes, while also identifying limitations in study design and the need for more rigorous validation.2

That is closer to the medicine I recognize: useful capability, uneven evidence, unfinished work.

AI is not one intervention. A system that shortens image interpretation, an algorithm that predicts deterioration, and software that encourages medication adherence are doing different jobs. Their benefits and risks should not be collapsed into one verdict about “AI in medicine.”

The value of technology lies partly in what it can see that we cannot. Its limits begin when seeing is mistaken for deciding.

What draws me to these technologies is not the prospect of removing people from medicine.

It is the possibility of extending what people can notice.

A clinician may see the patient four times a year. A device may reveal the trajectory between those visits. An algorithm may surface an abnormality buried in information that nobody had time to examine together. Well-designed systems may also take repetitive work away from clinicians and return some attention to the parts of medicine that require a human being in the room.

I want those capabilities to improve.

But there is a boundary worth protecting.

The value of technology lies partly in what it can see that we cannot. Its limits begin when seeing is mistaken for deciding.

An accurate measurement can still answer the wrong question

A patient asking, “Is my pulse fast?” can now often ask a watch.

But medicine rarely ends with that answer.

Why is it fast? Does the rhythm matter? Does it explain the symptom? Should anything change because of it?

Those are different questions.

A device may measure one variable well without establishing why that variable changed. An irregular-rhythm notification does not know the patient’s complete history, current symptoms, recent infection, thyroid status, medications, or what the conventional ECG showed six months earlier.

The same problem appears with prediction.

A model may estimate risk well in the population in which it was developed and still be less informative for the particular person sitting across from us. An AI system can recognize a pattern while missing the one piece of context that would change its interpretation.

At the bedside, I separate three things.

First, did the tool measure something credibly?

Then, does that measurement mean something clinically?

Only after those questions do we reach the one patients actually care about: does knowing this improve what we should do?

Technology can often answer the first question quickly.

The last one may require symptoms, history, competing risks, other evidence, uncertainty, and the preferences of the person whose body is being measured.

A beautifully measured number can still answer a question nobody needed to ask.

More monitoring creates more interpretation

I was trained in a medicine where missing information could be dangerous.

That makes abundance feel like progress.

Often it is.

For decades, much of physiology between appointments simply disappeared from view. We had the blood pressure measured in the clinic, the ECG captured during symptoms if timing was kind, and laboratory values taken at intervals. The remaining hours belonged largely to inference.

Digital medicine has changed that scarcity.

A physiological stream can now continue while a person sleeps, works, walks, eats, exercises, becomes ill, and recovers.

But biological variation did not disappear when monitoring improved.

Heart rate still responds to sleep, hydration, infection, exercise, pain, temperature, emotion, medication, and posture. Blood pressure still moves through the day. Glucose still changes with meals, activity, illness, stress, and sleep.

Dense measurement makes meaningful patterns easier to discover.

It also makes ordinary deviations easier to find.

The challenge therefore changes from finding information to deciding which information deserves attention.

A 2025 systematic review examined 116 randomized trials of device-based remote monitoring. Across the comparisons reviewed, hospital-service use was descriptively lower in many remote-monitoring groups, but the authors’ examination of program features did not show that more frequent provider assessment of data consistently distinguished better-performing approaches. Daily measurement and transmission were common, yet higher frequency itself did not translate straightforwardly into greater reductions in hospital use. The authors emphasized the wider care process, including support, staffing, integration, and redesign, rather than technology alone.3

Because the trials were highly heterogeneous and these component comparisons were descriptive, we should not conclude that less monitoring is better.

What I take from the evidence is narrower.

More frequent measurement is not automatically more useful measurement.

A thousand observations create a thousand opportunities to notice something.

They also create more material that someone, or something, has to interpret.

Perhaps the useful unit in digital medicine is not the measurement itself.

It is the decision the measurement improves.

The tool changes the person using it

Technology does not simply observe us.

We begin observing ourselves through it.

That can be helpful. A glucose trace may turn an abstract dietary discussion into something visible in a person’s own afternoon. A blood-pressure record can replace memory with a pattern. Activity tracking can reveal how quietly movement has disappeared from a week.

Measurement can create agency.

It can also redirect attention.

In a 2024 study of 172 people with atrial fibrillation, 83 used wearable devices. Wearable users reported more symptom monitoring and preoccupation, more concerns about their AF treatment, and greater AF-related healthcare use than nonusers. Among wearable users, 20 percent reported intense anxiety in response to irregular-rhythm notifications.4

The study was observational. It cannot tell us that the devices caused those differences. People who are already more vigilant or concerned about their arrhythmia may also be more likely to own and intensively use a wearable.

The same study contained an apparently contradictory finding.

Nearly two thirds of wearable users said the device made them feel safer.4

I think both results belong in the same paragraph.

The same device may give one patient confidence and another more reasons to worry. It may encourage useful self-management or slowly train attention toward every fluctuation.

A health tool therefore has effects beyond the accuracy of its sensor.

It can change behavior.

That matters because the response to information is part of its clinical consequence.

An accurate reading can still lead to an unnecessary call.

A reassuring screen can also become a reason to discount a symptom that deserves attention.

The device may be working perfectly in both cases.

The difficulty lies in the authority we give its answer.

A prediction is not a command

Artificial intelligence moves the problem one step further.

A sensor measures.

AI may classify, predict, prioritize, summarize, or recommend.

A risk estimate becomes useful only when we understand enough about the population behind it, the time horizon it describes, the information the model considered, and what relevant context may be missing.

Then we have to consider the consequences if the prediction is wrong.

FDA’s January 2026 guidance on clinical decision-support software draws a revealing line. For certain decision-support functions to fall outside regulation as medical devices, clinicians must be able to independently review the basis of a recommendation rather than rely primarily on the software’s conclusion.5

FDA is drawing a regulatory line, but the clinical principle matters too.

A recommendation deserves more confidence when its basis can be questioned. We need to know what evidence produced it, how well the relevant population was represented, what limitations are known, and what information would make us distrust the result.

Without that ability, probability can begin to feel like instruction.

It is not.

A prediction can narrow uncertainty. It cannot decide what uncertainty a particular person should accept.

An algorithm may estimate the probability of deterioration or benefit from another test. It cannot decide how one person weighs the consequence of a missed diagnosis against the burden of another invasive procedure.

That part of medicine does not disappear because the prediction became more accurate.

In some ways, the better the prediction becomes, the more important the distinction becomes.

Putting a human in the loop is not enough

The familiar reassurance about medical AI is that a human will remain involved.

That sounds sensible.

It may not be sufficient.

A clinician can technically be “in the loop” while lacking enough time to reconsider an output, enough knowledge to understand its limitations, enough authority to override it, or any practical ability to stop what the system has already set in motion.

A 2026 npj Digital Medicine commentary describes four requirements for meaningful oversight: epistemic capacity, cognitive space, decisional authority, and effective ability to intervene. Human presence alone, the authors argue, does not create meaningful oversight.6

There is another reason.

People do not receive automated advice neutrally.

The literature has long described automation bias, the tendency under certain conditions to over-rely on automated decision support. A systematic review found that such systems can improve performance while also creating new kinds of error when users accept faulty advice or reduce independent checking. Workload, time pressure, confidence in the system, and the way advice is presented can influence that behavior.7

A 2017 simulated electronic-prescribing experiment made the problem concrete. One hundred and twenty senior medical students worked through prescribing scenarios with correct, incorrect, or absent decision support. Correct support reduced omission errors. Incorrect support increased them, and false recommendations were accepted surprisingly often.8

A simulation with medical students is not the same as clinical practice with experienced physicians.

But the mechanism deserves attention.

A system that is usually right can change the psychological cost of disagreement. Under pressure, its recommendation may gradually become the starting assumption rather than one piece of evidence.

An algorithm does not have to replace judgment to weaken it. Sometimes it only has to become the answer we stop challenging.

Surgeons understand a version of this problem without using the language of AI.

Being present in an operating room does not mean someone is capable of preventing an error. Responsibility requires enough understanding to recognize what is happening and enough authority to intervene before the consequence is fixed.

Digital medicine deserves the same seriousness.

Human oversight should mean that the human can still think, question, and change what happens next.

Technology does not reach everyone equally, or know everyone equally well

One of digital medicine’s strongest promises is simple: bring care closer.

For many patients, it does.

Remote monitoring can spare difficult travel. Home measurements can extend care to people who live far from specialists. Digital tools can make regular observation possible when repeated conventional visits would be burdensome.

But “remote” does not automatically mean “accessible.”

From the clinic, sending a measurement may look almost effortless. At home it can depend on a compatible device, stable connectivity, sufficient vision and dexterity, language the person understands, enough digital confidence to navigate the interface, and help when something stops working.

A 2025 equity analysis examined 119 reports of remote-monitoring programs for chronic disease against eleven equity-related parameters. Reporting was inconsistent and incomplete. Only 7 percent reported inclusion of people with varying levels of digital literacy, and 4 percent reported inclusion of people with physical or mental disabilities.9 The authors appropriately distinguished non-reporting from proven exclusion, but their findings challenge the assumption that remote care becomes equitable simply because it moves outside a clinic.

Access is only one layer.

Digital inequality can begin before a patient ever touches the device.

Algorithms learn from data. If some populations are poorly represented in the data used to develop or test a system, access to the same technology does not guarantee access to equally reliable inference.

FDA, Health Canada, and the UK’s Medicines and Healthcare products Regulatory Agency make this explicit in their transparency principles for machine-learning-enabled medical devices. Developers are encouraged to identify gaps in data characterization, including patient populations not well represented in training or clinical datasets and therefore potentially at greater risk of bias.10

The interface can look equally sophisticated for everyone while the confidence we should place in its output differs.

There is another inequality that interests me because it is easier to miss.

A patient with a recent smartphone, reliable internet, several connected devices, confidence with technology, and ready access to clinicians may generate thousands of data points.

Another patient may own no wearable, struggle to reach routine care, speak a language poorly supported by the system, or have no stable connection through which to transmit anything.

The first person becomes increasingly visible to medicine.

The second may remain quiet in the data.

If we are careless, the people easiest to measure may become easier to notice, while some of the people already hardest to reach become quieter still.

That is not a reason to collect less useful information.

It is a reason not to confuse abundance of data with importance of need.

Technology can reduce one kind of distance while creating or exposing another.

Equity therefore belongs inside our judgment about whether a technology works, not in a separate conversation after the system has already been built.

What deserves to change the decision?

The patient’s telephone is still on my desk.

There are many numbers.

The task is not to dismiss them.

It is not to honor all of them equally either.

One trend makes me return to a part of his history.

Another is interesting but does not alter what we know.

A third may deserve confirmation by a different method before it is allowed to carry much clinical weight.

Several change nothing.

I have become less interested in asking whether I “trust” a device in the abstract.

Trust is too large a verdict.

A better question is how much weight this particular piece of information has earned in this particular decision.

Credibility matters, but credibility is not enough. The result has to survive context. It has to fit, or meaningfully challenge, the symptoms, history, and other evidence. Even then, information may add little if no consequential decision depends on it.

The health system faces the same problem at scale.

A 2025 report describing Stanford Medicine’s work on patient-generated health data found that clinical integration required much more than accepting incoming data. Their workgroup emphasized setting expectations about how information would be used, preparing staffing and workflows, and establishing how outlying values would be handled.11 The authors were describing one academic health system, not a universal model, but the lesson is recognizable: transmitting data and using data are different achievements.

This is where Tools & Decisions meets Medicine’s Blind Spot.

Medicine’s ability to see is expanding rapidly.

We can follow physiology outside the clinic, detect patterns earlier, and place extraordinary computational capacity beside human judgment.

I want that progress.

The blind spot appears when visibility is mistaken for understanding, precision for relevance, or prediction for instruction. It also appears when the easiest patient to measure becomes the easiest patient to notice, or when a recommendation acquires more authority than its evidence has earned.

I think again of the patient and the phone between us.

He arrived with more information than a patient could have brought into my office thirty years ago.

That is progress too.

But the useful outcome of the consultation is not that we have discussed every graph.

We have done something harder.

We have decided which ones deserve consequences.

Technology will keep extending medicine’s sight.

Our responsibility is to decide what should follow from what it allows us to see.

References

1. U.S. Government Accountability Office. Wearable Technologies: Potential Benefits and Challenges in Clinical Decision-Making. GAO-26-107847. Washington, DC: U.S. Government Accountability Office; August 6, 2026.

2. Lin YE, Yang SM, Huang CJ, et al. Impact of artificial intelligence on cardiovascular workflow, engagement, and outcomes: a systematic review. npj Digital Medicine. 2026;9:536. doi:10.1038/s41746-026-02690-7.

3. Jansen AJS, Peters GM, Kooij L, Doggen CJM, van Harten WH. Device based monitoring in digital care and its impact on hospital service use. npj Digital Medicine. 2025;8:16. doi:10.1038/s41746-024-01427-8.

4. Rosman L, Lampert R, Zhuo S, et al. Wearable devices, health care use, and psychological well-being in patients with atrial fibrillation. Journal of the American Heart Association. 2024;13(15):e033750. doi:10.1161/JAHA.123.033750.

5. U.S. Food and Drug Administration. Clinical Decision Support Software: Guidance for Industry and Food and Drug Administration Staff. January 2026.

6. van de Sande D, Economou-Zavlanos N, van Genderen ME. Meaningful oversight of medical AI beyond human in the loop. npj Digital Medicine. 2026;9:569. doi:10.1038/s41746-026-02971-1.

7. Goddard K, Roudsari A, Wyatt JC. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association. 2012;19(1):121-127. doi:10.1136/amiajnl-2011-000089.

8. Lyell D, Magrabi F, Raban MZ, et al. Automation bias in electronic prescribing. BMC Medical Informatics and Decision Making. 2017;17:28. doi:10.1186/s12911-017-0425-5.

9. Abejirinde IOO, Kishimoto V, Pfisterer KJ, et al. An equity analysis of remote patient monitoring programs unveils assumptions on digital health equity. npj Digital Medicine. 2025;8:320. doi:10.1038/s41746-025-01731-x.

10. U.S. Food and Drug Administration, Health Canada, Medicines and Healthcare products Regulatory Agency. Transparency for Machine Learning-Enabled Medical Devices: Guiding Principles. June 2024.

11. Griffin AC, Moyer MF, Anoshiravani A, Hornsey S, Sharp CD. A sociotechnical approach to defining clinical responsibilities for patient-generated health data. npj Digital Medicine. 2025;8:270. doi:10.1038/s41746-025-01680-5.

Author’s note: The opening scene is a composite drawn from recurring clinical situations. It does not represent a single identifiable patient.

About the author
Mustafa Kemal Calik, MD, is a cardiovascular surgeon and digital health consultant. He writes about the space between medical capability and ordinary life; how care, technology, relationships, and daily conditions shape what happens after the clinical decision is made.

mustafa-kemal-calik-md-cardiovascular-surgeon