ARTICLE 11 │ PRACTITIONER VOICES IN ETHICAL AI │ PART ONE │ JULY 2026
What We Cannot Answer Alone
Healthcare has documented its own inequities for decades. Now AI is learning from that record. Before we scale, we need practitioners at the table.
By Constantine Rhaich’al │ Co-Founder & CEO, Synod IntelliCare Inc.
A health record is a kind of memory. It remembers what was given: the scan ordered, the analgesic dispensed, the follow-up appointment booked. What it does not remember, and was never designed to remember, is what was needed and withheld. That asymmetry sat harmlessly in filing cabinets for fifty years. It does not sit harmlessly now.
Canadian health systems have documented their own inequities for decades. The Canadian Institute for Health Information (CIHI) has tracked persistent gaps in outcomes by income, geography, immigration status, and race across our health systems (CIHI, 2024). The same findings appear year after year, report after report, with a consistency that should trouble anyone who believes measurement alone produces change. The problem was never that we did not know. The problem was that knowing did not oblige anyone to act. Institutional knowledge of inequity, without institutional action, is a governance failure, and it leaves a paper trail. The machines are now reading it.
Obermeyer et al. (2019) showed the mechanism plainly in Science. A widely used commercial algorithm systematically underestimated illness severity in Black patients. No one wrote a discriminatory rule. The model used cost as a proxy for need, and Black patients at equivalent levels of illness had historically received less care, so the record showed less cost, and the model concluded less need. The arithmetic was objective. The inheritance was the truth. This is no longer a question about a handful of pilots: Chang et al. (2025) report that 71% of U.S. hospitals used predictive artificial intelligence integrated with the electronic health record in 2024. The inheritance is already inside the workflow.
The thing I believe is most fragile here is not the model. It is the clinician's judgment. A nurse who has stood in a triage bay for fifteen years carries knowledge no dataset contains: the particular quality of a patient's stillness, the way a family member's account diverges from the chart, the sense that something is wrong before the vitals agree. That instinct has always been the last line of defense when a system's assumptions fail a patient. Deference to decision support is not laziness or malice; it is historical and rational. The tool is right most of the time, it is faster, and it never gets tired at hour eleven of a twelve-hour shift. But the times it is wrong are not randomly distributed. They cluster on exactly the patients the record under-served in the first place. And institutions build asymmetric friction around disagreement: agreeing with the model costs nothing, while overriding it is a deviation you may later have to justify. Under that asymmetry, a bias an empowered practitioner could have named out loud becomes the validated output of a trusted system.
I want to be careful about what I claim here, because I am not a clinician. I have never triaged a patient, assessed anyone's pain, or carried the weight of a decision made at three in the morning with incomplete information. I built this company in the hours after my daughter's bedtime, out of a conviction that fairness in these systems has to be measurable or it is only sentiment. I hold that conviction still, and I hold it alongside a limit I cannot argue my way past. Synod has built detection technology. We can surface performance differences across demographic groups in a model's outputs. What we cannot determine from outside a care setting is which of those differences is clinically meaningful. A few points of difference in a model's discrimination may be statistical noise in one department and a missed presentation in another. That threshold is not a mathematical question. It is a clinical and ethical one, and it belongs to the people who will live with the consequences of getting it wrong.
It would be easier, in some ways, if health executives disagreed with us. Disagreement is a persuasion problem, and persuasion problems have known solutions. What we found instead is harder. When we surveyed health system decision-makers this spring, sixty percent already rated the urgency of AI bias as high or immediate (Synod IntelliCare, 2026). They are not waiting to be convinced. They are waiting for something to do. Across the institutions we speak with, oversight committees are forming, principles documents are being drafted and revised, and someone has been named the AI lead, usually alongside three other portfolios. What almost none of them have is an operational mechanism to answer a narrow, uncomfortable question: are the AI tools already running inside this environment producing different outcomes for different patient populations?
When we press on why that gap persists, the answer is almost always practical rather than philosophical. Fairness auditing sounds like a systems project. Systems projects mean procurement cycles, security review, interface work, an integration queue already full for eighteen months. The institution concludes, reasonably, that it cannot generate evidence until the technical work is finished. That assumption is wrong, and dismantling it may be the most useful thing this article can do. A retrospective fairness audit of de-identified historical data does not require live integration with an electronic health record. It requires a data extract, a governance agreement, and a defined analytic protocol. We have started calling this the “sidecar” approach, because it runs alongside the vehicle rather than waiting to be built into the engine. What it decouples matters more than the method itself: it separates governance readiness from IT capacity. Those two things have been fused in institutional planning for years, and the fusion has made a great many organizations feel that acting responsibly is something they will be able to afford later. Trust-building can begin before the technical integration is complete.
The timeline is not being set by vendors. The EU Artificial Intelligence Act classifies many systems used for diagnosis, clinical decision support, treatment recommendations, triage, and patient monitoring as high-risk, and attaches to that classification a continuing obligation, risk management across the lifecycle, post-market monitoring, and documented bias mitigation (European Parliament, 2024). In the United States, the 2024 Section 1557 final rule requires covered entities to make reasonable efforts to identify and mitigate discrimination arising from patient care decision support tools, including AI-based ones (U.S. Department of Health and Human Services, 2024). Colorado's AI Act adds risk management programs, annual impact assessments, and documented mitigation for high-risk healthcare applications. Canada relies on a mix of federal frameworks, AI4H, and provincial obligations attaching to privacy and automated decision systems, pushing hospitals and vendors toward documented AI governance. Quebec’s Law 25 mandates duty-to-inform requirements, and other provinces follow their respective privacy laws, like Ontario’s PHIPA. The landscape is rapidly evolving; technology often advances faster than legislation, though liability trends are aligned. Early AI malpractice analyses focus less on whether a model was accurate in the abstract and more on how it shaped a physician’s judgment at the point of decision, which means algorithmic bias, inadequate testing, and absent oversight can distribute exposure across developers, institutions, and clinicians simultaneously. No party in that chain can fully insulate itself by pointing at another.
Everything I have written so far is a thesis. Healthcare AI systems inherit the inequities of the data they are trained on; those inequities surface as clinical harm; and the harm is measurable, therefore preventable. I believe that. Belief is not the same thing as proof, and I want to be precise about which one we currently hold. The external support is real, Obermeyer, and the growing literature on demographic skew in language model outputs, but those are other people's findings. They justify the question. They do not answer it for Canadian clinical contexts, and they do not validate anything we have built. So the honest description of where we are is this: we are building the machinery to test our own thesis, in public, with people who can tell us if we are wrong.
Our Connected Minds research, “Towards Fair and Reliable Large Language Models in Healthcare”, is now in execution with Prof. Ines Arous of York University's Lassonde School of Engineering and Prof. Afaf Taik of the Université de Sherbrooke. The problem it addresses is not speculative. Language models are already inside Canadian clinical workflows, drafting documentation, summarizing records, and mediating patient communication. Physicians in Ontario emergency departments report using clinical scribes that document encounters automatically, and hospitals are exploring conversational agents for patient intake. Deployment has outrun evaluation. There is currently no structured bias evaluation framework tailored to Canadian clinical contexts, and the work that exists relies on ad hoc, non-transparent methods that cannot be reproduced or compared. The research project aims to design one, in the paediatric domain to start, specifically: evaluating how demographic variables - race, gender, socioeconomic status - shift the content of generated clinical answers, and investigating explainability techniques to locate why and where bias enters model outputs rather than simply confirming that it did. There is a constraint I would rather name than discover in review. The work requires paediatric clinical text, and existing datasets overwhelmingly cover adults. That gap is not an inconvenience in our project plan. It is a structural feature of the field, and it is one of the reasons institutional partnership is a prerequisite rather than a preference.
The second strand sits with Sheridan College's Centre for Applied AI: an interactive dashboard that lets healthcare stakeholders monitor, interpret, and evaluate AI-driven clinical decision systems transparently. It responds to a problem our partners identified, which is that current tools give clinicians and administrators almost no real-time visibility into how a deployed system is behaving. Without that visibility, errors, biases, and inconsistencies are found retrospectively, if at all. Both of these projects converge on the same engineering problem, and it is not a detection problem. Our Data Diversity and Fairness Auditor (DDFA) works, it identifies intersectional bias in clinical datasets, and its output is excellent for data scientists, policy leads, and risk managers. For clinicians, it is close to unusable. That is a translation problem. The Clinical Relevance Dashboard is our answer: a generative engine that converts raw fairness metrics into a Composite Fairness Score (CFS), and then converts that score into structured, evidence-based clinical narratives, measurements grounded in medical context, calibrated and reproducible enough to withstand academic and institutional scrutiny. A fairness metric a clinician cannot act on is not a fairness intervention. It is a number without context.

We could have shipped and asked forgiveness. Plenty of companies in this market have. Building the research foundation first is slower, more expensive, and harder to explain to investors, and it is the point: a company arguing that healthcare AI should be validated before deployment cannot itself deploy before validation. That is not virtue. It is the minimum coherence requirement of our own argument. So here is the balance sheet. We are pre-revenue. The research is in execution, not complete. No finding has been peer-reviewed. The Composite Fairness Score (CFS) does not yet exist in a form anyone outside our team has stress-tested. Our thesis is being tested, not confirmed, and it is entirely possible that some of what we believe will not survive contact with the data.
What we are asking practitioners for falls into four areas, and none of them is endorsement. The first is governance: help us understand how fairness auditing can embed into the oversight infrastructure your institution is already building, rather than asking you to build a parallel one. The second is regulatory and compliance rigor: we prioritize PHIPA/HIPPA and PIPEDA from day one, alongside alignment with Health Canada's AI4H direction and the OCAP and CARE frameworks, and we want to be told where our reading of those obligations is thin. The third is clinical impact and outcomes: tell us what counts as a meaningful disparity where you work, and what size of gap would actually change your practice. We are deliberately not prescribing success metrics. We would rather understand the indicators your own committees hold you to, around patient safety, equity, and the Quadruple Aim, so that the evidence we build together is legible to them rather than to our roadmap. The fourth is capacity-building, and it is not an afterthought. A fairness metric no one on staff can interpret is not oversight; it is theatre with a dashboard. Our work with community health centres is built on co-design, so that tools fit the workflow that actually exists rather than the one described in a process map; on practical training in bias detection and health equity, so staff can read a fairness metric in the course of their day and know what it is telling them; and on sustainability, so the institution keeps an implementation toolkit and bias-aware protocols it can run independently after we are gone.
This article opens a four-part series, and the three that follow are not mine. Next month we will have a retired practitioner with more than twenty years of nursing and healthcare operations leadership. September, we will have a professor write on the social determinants that shape a record long before an algorithm ever reads it. Our co-founder and a practicing nurse of twenty years writes in October on education, workforce, and the internationally educated practitioners whose training the system has never fully valued. No one wanted to go first, so I was voluntold to lay out the vision and prep the ground.
So the invitation is specific. Tell us the ethical challenge you have already met in your own setting, the one you have not seen described accurately anywhere. Tell us what a meaningful disparity looks like where you work. And tell us what your institution already holds you accountable for, because a fairness measure that cannot live alongside your existing obligations will not survive contact with a real ward. We will publish what we learn, including the parts that complicate our own position. I would rather build something slowly with people who know what they are talking about than quickly with people who agree with me.
An Invitation
If you are a practitioner, educator, or health system leader with an ethical AI challenge you have not seen named properly, we want to hear it. That is the whole of the ask.
If your organization is building AI oversight infrastructure and wants to understand where fairness auditing fits, we would value a governance scoping conversation before any discussion of deployment.
If you want to know where your organization stands on ethical AI maturity right now, the Ethical AI Maturity Assessment takes less than fifteen minutes.
Take the Ethical AI Maturity Assessment: Ethical AI Maturity Assessment (EAMA)
Start a conversation: Contact Us
Standards do not arrive fully formed. They are built, by the people willing to take them seriously before they are required.
About the Author
Constantine Rhaich’al is Co-Founder and CEO of Synod IntelliCare Inc., a Toronto healthcare AI ethics company building fairness auditing infrastructure for clinical AI. He spent twenty years in enterprise healthcare before founding Synod, which he built in the quiet hours after his daughter’s bedtime. This article opens the Practitioner Voices in Ethical AI series.
References
Canadian Institute for Health Information. (2024). Health system performance and health equity in Canada. CIHI. https://www.cihi.ca
Chang, W., Owusu-Mensah, P., Everson, J., & Richwine, C. (2025). Hospital trends in the use, evaluation, and governance of predictive AI, 2023–2024 (Data Brief No. 80). Office of the Assistant Secretary for Technology Policy.
European Parliament. (2024). Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union.
Ghassemi, M., Oakden-Rayner, L., & Beam, A. L. (2021). The false hope of current approaches to explainable artificial intelligence in health care. The Lancet Digital Health, 3(11), e745–e750.
Obermeyer, Z., Powers, B., Vogeli, C., & Mullainathan, S. (2019). Dissecting racial bias in an algorithm used to manage the health of populations. Science, 366(6464), 447–453. https://doi.org/10.1126/science.aax2342
Omiye, J. A., Lester, J. C., Spichak, S., Rotemberg, V., & Daneshjou, R. (2023). Large language models propagate race-based medicine. npj Digital Medicine, 6, 195.
Synod IntelliCare. (2026). Willingness-to-Buy Survey: Clinician and health executive perspectives on AI fairness. Unpublished survey data.
U.S. Department of Health and Human Services. (2024). Nondiscrimination in health programs and activities: Final rule (Section 1557). Office for Civil Rights. https://www.hhs.gov
Zack, T., Lehman, E., Suzgun, M., Rodriguez, J. A., Celi, L. A., Gichoya, J., Jurafsky, D., Szolovits, P., Bates, D. W., Abdulnour, R.-E. E., Butte, A. J., & Alsentzer, E. (2024). Assessing the potential of GPT-4 to perpetuate racial and gender biases in health care. The Lancet Digital Health, 6(1), e12–e22.