Trusted information, accountable governance, and secure access as the basis for reliable healthcare AI
A hospital introduces an AI assistant to help clinicians prepare discharge summaries. The assistant can search the electronic health record, summarize notes, and assemble a medication history. During a demonstration, it produces a polished account of a patient’s admission. Yet its underlying sources include an unreconciled medication list, a preliminary laboratory result that was later corrected, and an old note copied into the current encounter. The prose is clear. The account is clinically misleading.
This hypothetical example illustrates why a healthcare AI initiative must begin with a data strategy. A model operates on the information supplied through its training, retrieval systems, application integrations, and immediate context. Errors can enter through any of these paths. A better model may recognize some inconsistencies, but an organization cannot depend on model intelligence to repair unidentified patients, missing records, incorrect units, obsolete policies, or inappropriate access.
Trusted data is necessary for accurate outputs, but it is insufficient by itself. A system can still retrieve the wrong evidence, misinterpret an accurate record, or generate an unsupported statement. The foundation therefore has to connect data quality with application design, clinical validation, security, and continuing oversight. Trust should mean that information has a known origin, an understood meaning, an appropriate level of quality for its intended use, and enforceable conditions governing access.
The strategy begins by defining the decision or workflow that AI will support. An assistant that searches hospital policies, a model that forecasts bed demand, and a system that flags clinical deterioration have different data requirements and consequences of failure. Each use case needs a documented purpose, intended users and population, required sources, acceptable freshness, evaluation criteria, and a named owner responsible for the resulting behavior. A clinical sponsor should establish which errors matter most and when the system must defer to a person.
That definition also separates three important data needs. Training data shapes a model’s learned behavior. Evaluation data establishes how it performs under specified conditions. Inference data supplies the information used when the deployed application answers a question or makes a prediction. Retrieval-augmented generation, or RAG, adds a searchable evidence collection to the inference process. Improving one of these data sources does not automatically improve the others: updating a policy repository changes retrieved evidence without necessarily changing model weights.
For a discharge assistant, the organization might require a reconciled medication source, encounter-specific laboratory results with their status, and clinician review before a draft becomes part of the signed record. For an administrative forecast, a daily feed may be sufficient. These are choices the organization must make explicitly. The important question is whether the available data can support the intended action within an acceptable margin of uncertainty.
Before building integrations, the organization needs an inventory of where relevant information originates and how it moves. Healthcare data may be distributed across an electronic health record, laboratory and radiology systems, pharmacy applications, claims platforms, scheduling systems, patient portals, scanned documents, and external exchanges. An inventory should identify each source’s owner, identifiers, format, update cadence, known defects, retention requirements, sensitivity, and approved uses. It should also identify copies in warehouses, spreadsheets, exports, vendor platforms, and analytical environments.
Fragmentation creates more than an integration problem. Two systems may describe the same event differently, update at different times, or retain conflicting versions. A pharmacy dispensing record, a medication order, and a patient’s report of what they actually take answer different questions. Combining them into one undifferentiated medication field loses meaning. A strategy should define how these sources relate and how unresolved discrepancies are presented, rather than declaring a single system authoritative for every question.
Patient identity is a particularly important boundary. Records cannot be combined safely merely because names and birth dates look similar. An enterprise patient identity service needs documented matching rules, controls for duplicates and mistaken merges, and review of ambiguous cases. False matches can expose one patient’s information to another patient’s care team and contaminate clinical inputs. Historical datasets also need a way to account for subsequent identity corrections.
Data quality then becomes a set of measurable, use-specific expectations. The following dimensions provide a practical starting point; they are an implementation framework, not a universal healthcare certification standard.
| Dimension | Question to resolve | Example of a useful control |
|---|---|---|
| Accuracy | Does the value reflect the underlying event? | Review sampled records against source documentation. |
| Completeness | Is the information required for this task present? | Measure missing critical fields by site, workflow, and population. |
| Consistency | Do related records agree, or explain their differences? | Detect conflicting identifiers, units, and medication statuses. |
| Timeliness | Is the information current enough for this decision? | Monitor source-to-application delay and enforce use-specific freshness limits. |
| Validity | Does the record follow the expected structure and meaning? | Validate schemas, terminology versions, units, and result status. |
| Provenance | Can its origin and transformations be reconstructed? | Preserve source identifiers, versions, and transformation history. |
| Relevance | Does the dataset represent the intended setting and task? | Evaluate coverage of the patients, sites, and workflows the application will encounter. |
A quality program needs both automated checks and clinical interpretation. Pipelines can detect unexpected schema changes, orphaned identifiers, duplicate messages, invalid codes, missing units, and unusual distributions. They can also flag implausible values for investigation. However, an unusual result may describe a genuinely unusual patient. Automatically deleting outliers or filling every empty field can remove clinically important information and conceal uncertainty.
Missing information deserves explicit treatment. A missing allergy entry does not establish that the patient has no allergies. An absent laboratory result may mean that the test was never ordered, remains pending, failed to transmit, or exists elsewhere. These distinctions can affect both clinical interpretation and model behavior. The data representation should preserve the reason for absence where available, and the application should expose relevant gaps instead of silently converting them into negative findings.
Quality also depends on context. A diagnosis may be confirmed, provisional, ruled out, historical, or part of a family history. A symptom can be present or explicitly denied. A medication may be prescribed, dispensed, administered, discontinued, or reported by the patient. A numerical observation needs its units, timing, measurement method, and relevant encounter or specimen information. Removing these distinctions makes a dataset easier to process while making it easier to misunderstand.
Interoperability standards help preserve structure. HL7 FHIR, for example, provides resource models in which observations can carry result status, timing, units, reference ranges, and reasons for missing data. Conditions can carry clinical and verification status. Implementers still need to preserve and interpret these fields correctly; an exchange format cannot establish the truth of the underlying clinical assertion. (HL7 FHIR Observation; HL7 FHIR Condition.)
A terminology strategy should likewise maintain explicit mappings among local codes and shared vocabularies, with versioned definitions and clinical review. Laboratory tests, diagnoses, medications, and units have different coding needs. Mapping should retain the original value and disclose unresolved equivalences. A standardized field is only useful if the transformation preserves the distinction that matters to the application.
Outdated data requires more careful handling than simply deleting old records. A remote surgical history may remain important, while yesterday’s medication list may already be obsolete. Freshness depends on the kind of information and the intended decision. The strategy should distinguish a historically valid observation from an assertion that is supposed to describe the patient’s current state.
At least two timelines often matter: when an event occurred and when the application could actually know about it. A specimen might be collected in the morning, its result released later, and a correction published the following day. Recording only specimen time obscures those changes. Where relevant, the pipeline should retain event time, source release time, ingestion time, and version history. Predictions made retrospectively must use the versions available at the decision time; otherwise, later results and corrections can leak into the evaluation.
The same principle applies to reference information. A clinical guideline, formulary entry, or hospital policy should have an identifiable owner, version, effective date, review status, and applicability. A RAG system needs a defined process for replacing superseded documents and invalidating derived indexes or cached answers. The newest document is not automatically the appropriate one for every historical question, but an expired policy should not quietly support advice about current operations.
Governance makes these expectations sustainable. Executive leadership should establish the organization’s priorities, funding, and risk tolerance. Clinical and operational leaders should own the meaning and approved uses of their data domains. Data stewards should maintain definitions, quality rules, and issue resolution. Engineering teams should implement reliable pipelines; security and privacy teams should establish protection and access requirements. A governance forum should resolve disputes and escalation decisions without becoming a substitute for day-to-day ownership.
Ownership in this setting means accountability for stewardship. It does not confer unrestricted authority to reuse patient information. A useful operating arrangement combines central standards for identity, security, metadata, and documentation with domain expertise in nursing, pharmacy, laboratory medicine, finance, and other areas. This allows decisions about clinical meaning to remain close to the people who understand how the records are produced.
Each approved dataset should have a documented agreement between its producers and consumers. Sometimes called a data contract, this should describe its schema, meanings, permissible purposes, freshness expectations, quality checks, owner, access conditions, and change notification process. If a source begins supplying a different unit or changes the definition of an encounter category, consumers should learn about that change before it silently alters model behavior.
Organizations must also examine how their data reflects clinical practice and access to care. Billing codes are shaped by reimbursement requirements. Test ordering reflects clinician decisions and resource availability. Documentation reflects time pressure, templates, and local habits. A record of service utilization is therefore not a neutral measure of underlying medical need.
A well-known study by Obermeyer and colleagues demonstrated this problem in a population health algorithm. The algorithm used healthcare cost as a proxy for health need, producing racial bias because spending did not represent need equally across groups. The finding illustrates why selecting the prediction target is itself a data governance decision. A model can accurately predict a proxy while failing to serve the intended clinical purpose. (Obermeyer et al., Science (2019).)
Evaluation should examine relevant differences across sites, age groups, languages, clinical populations, and other appropriate characteristics. Removing a demographic field does not establish fairness, because related patterns may remain elsewhere in the data. Conversely, collecting sensitive attributes solely because they might be useful requires a defined purpose and appropriate safeguards. The organization needs evidence that its application works for the population it intends to serve.
The technical architecture should make approved data use concrete. Source records can flow through validation and reconciliation into governed datasets, with separately controlled environments for development, evaluation, and production inference. A catalog should record definitions, owners, sensitivity, limitations, and lineage. Reusable datasets should have explicit intended uses rather than becoming a general pool from which every AI project draws indiscriminately.
Lineage must extend beyond the original database. For a training run, retain the dataset version, inclusion criteria, transformations, terminology mappings, code version, and model configuration. For a generated clinical draft, retain enough controlled evidence to identify the source versions and application configuration that produced it. FHIR’s Provenance resource supports recording the activities and agents involved in creating or revising resource versions; its AuditEvent resource supports a related record of system activity and access. These mechanisms support accountability, but require implementation and dependable collection. (HL7 FHIR Provenance.)
Live access requires a different control from approval of a historical training snapshot. A clinician’s ability to view a chart in the EHR should not automatically become unrestricted access through a conversational interface. The AI application’s effective permissions should reflect the authenticated user, application identity, permitted purpose, patient or encounter context, and applicable restrictions.
For RAG, authorization must be enforced before protected content reaches the model. Search and retrieval should select only records the requester is allowed to use. Enforce those decisions again where necessary when content is fetched, passed to tools, or released to a destination. Asking the model to avoid unauthorized material after supplying it creates an unnecessary disclosure path. Search results, snippets, and document titles can themselves reveal sensitive information.
Derived data needs the same attention. Embeddings, cached answers, conversation histories, traces, and evaluation examples may retain sensitive content or information about it. Cache design must prevent one patient’s answer from being reused for another patient or one user’s permissions from being inherited by another user. Permission changes and record corrections should propagate to affected retrieval indexes and caches within defined time limits.
An agent that can invoke tools adds a further permission boundary. Its service identity should have only the capabilities required for its approved workflow, with access constrained by the requesting user’s authorization where it acts on that user’s behalf. Reading a chart, drafting a note, signing documentation, sending a message, and submitting an order are distinct operations. Each needs explicit authorization and appropriate validation. Critical changes should pass through established clinical approval workflows.
FHIR helps exchange healthcare information, but it does not provide a complete security system. HL7’s security guidance assigns implementers responsibility for authentication, authorization, secure communications, auditing, and applicable legal requirements. A FHIR endpoint therefore needs an enforceable security architecture around it, including protection against overly broad searches and excessive extraction. (HL7 FHIR security guidance.)
HIPAA compliance must be evaluated around the actual use of information. The Privacy Rule permits specified uses and disclosures, including treatment, payment, and healthcare operations, subject to its conditions. That does not establish that every proposed AI training project, commercial reuse, or vendor improvement program fits one of those purposes. Document the permitted basis for each use and disclosure, including movement to external services. (HHS summary of the HIPAA Privacy Rule.)
Research may require individual authorization or an applicable alternative, such as a documented waiver approved by an Institutional Review Board or Privacy Board. The organization should determine whether an activity is research, healthcare operations, or another use based on its actual purpose. Separately, certain substance use disorder records may be subject to 42 CFR Part 2. These classifications should influence access and processing rules. (HHS HIPAA and research guidance; HHS 42 CFR Part 2 fact sheet.)
The minimum-necessary standard also needs precise interpretation. It generally requires reasonable limits on uses, disclosures, and requests for PHI to accomplish their purpose. Among its exceptions are disclosures to, or requests by, healthcare providers for treatment purposes. That exception does not make all internal access or every AI-related use exempt. Internal access policies should identify the workforce roles, information categories, and conditions needed to perform their duties. Purpose-specific access can support care while limiting unnecessary exposure. (HHS minimum-necessary guidance.)
Vendor relationships deserve review across the full service chain. An AI provider that creates, receives, maintains, or transmits PHI on behalf of a covered entity may be a business associate. When that relationship applies, the required business associate agreement must be established, and relevant subcontractor arrangements must also be addressed. HHS guidance explicitly includes a third-party AI chatbot handling patient PHI among its business associate examples. An agreement must constrain permitted uses and disclosures; signing one does not make an otherwise impermissible use acceptable. (HHS business associate guidance.)
Cloud architecture does not eliminate these obligations. HHS explains that a cloud service provider maintaining encrypted ePHI can be a business associate even when it lacks the decryption key. Responsibility for particular controls depends on the arrangement, and the parties need to understand and document their respective responsibilities. (HHS HIPAA and cloud computing guidance.)
As a practical procurement requirement, review the exact services and configurations covered by contractual commitments. Determine whether prompts, outputs, logs, support access, and subcontractors are included; whether customer information can be used for training; how long copies persist; and how incidents, return, deletion, and termination are handled. A restriction on vendor training is a valuable protection, but it addresses one issue within a broader data-use arrangement.
The Security Rule requires safeguards for ePHI across administrative, physical, and technical domains to preserve confidentiality, integrity, and availability. Regulated organizations must assess risks and vulnerabilities thoroughly and accurately, then manage the risks identified. The assessment needs to account for the actual AI environment, including new repositories, external processors, interfaces, and copies of sensitive information. (HHS summary of the HIPAA Security Rule.)
Technical measures should follow that analysis and the applicable requirements. Useful engineering practices include strong authentication, narrowly scoped service credentials, encryption, key management, network segmentation, protected audit records, secure development, vulnerability management, and tested backup and recovery. An organization should distinguish a recommended architecture from a claim that every listed technique is a universal HIPAA mandate. The controls must also support clinical availability and a safe fallback when the AI service fails.
Marketing terminology provides little assurance by itself. HHS’s Office for Civil Rights does not endorse or certify particular technologies or products. A vendor’s description of a service as “HIPAA compliant” therefore needs evidence about the contract, configuration, safeguards, and actual deployment. (HHS on technology endorsement and certification.)
De-identification requires similar care. HIPAA recognizes Safe Harbor and Expert Determination methods. Removing names, masking a few fields, or replacing identifiers with tokens does not by itself establish that a dataset meets either method. Dates, rare combinations of attributes, narrative text, and other identifying details require consideration. A limited data set remains PHI and has separate requirements for permitted purposes and a data use agreement. (HHS de-identification guidance; HHS summary of the HIPAA Privacy Rule.)
For development, appropriately constructed synthetic or de-identified data can reduce unnecessary exposure. However, synthetic data should be assessed for leakage and clinical usefulness, and sensitive intermediate material should remain controlled. Treat embeddings, trained artifacts, and transformed datasets as potentially sensitive until their information content and intended release have been evaluated. Privacy protection and preservation of clinical meaning must be assessed together.
Security also protects the trustworthiness of the evidence. Unauthorized changes to reference material, dataset labels, or terminology mappings can alter outputs without changing the model. Retrieved documents may contain instructions attempting to redirect an assistant. Source approval, controlled publication, integrity checks, monitoring, and strict tool authorization help protect this boundary. Clinical text should be processed as evidence; its presence in a record should not grant it authority to change application permissions.
Audit design should balance traceability with exposure. A useful record can identify the requesting user, service identity, purpose, source versions, policy decision, model configuration, and resulting action. Full prompts and outputs may contain PHI, so unrestricted diagnostic logging can create another sensitive repository. Determine which details are needed, restrict their access, and establish retention periods appropriate to their purpose and applicable requirements.
Retention and correction must cover derived copies as well as source records. If a record is amended, superseded, or removed under an applicable policy, identify affected indexes, analytical datasets, cached responses, and development copies. Retention policies must account for medical record obligations, legal holds, and backup constraints. Training artifacts require particular planning because removing a source row does not necessarily remove information learned by a model.
Validation should test the entire application under realistic conditions. FDA-linked Good Machine Learning Practice principles emphasize representative datasets, independence of training and test data, clinically relevant testing, and monitoring deployed models. Where appropriate, evaluation should separate patients, use later time periods, and test additional sites. The split must match the intended deployment and prevent duplicate or related records from creating misleading performance estimates. (IMDRF Good Machine Learning Practice principles (2025).)
For a generative application, evaluate unsupported claims, omitted information, temporal accuracy, source attribution, and whether retrieved evidence actually supports the answer. Test failure conditions such as stale feeds, unavailable systems, ambiguous patient matches, access revocation, and incomplete documentation. A system should indicate when it lacks sufficient evidence rather than conceal a failed integration behind fluent prose.
Measures of success should extend beyond aggregate model accuracy. Track errors with potential clinical consequences, correction burden, workflow disruption, relevant subgroup performance, retrieval failures, and inappropriate disclosure attempts. A technically accurate assistant can still be unhelpful if reviewing its output takes more work than completing the task directly. Clinicians need a practical way to inspect sources, correct errors, and report recurring problems.
Corrections should feed a governed improvement process. Automatically treating edited AI output as ground truth can reproduce mistakes or introduce manipulated examples. Review feedback, identify whether the failure originated in the source, transformation, retrieval, model, or interface, and repair the responsible layer. Otherwise, repeated prompt changes can conceal a recurring upstream data problem.
NIST’s voluntary AI Risk Management Framework provides an organizing approach through Govern, Map, Measure, and Manage. Applied here, those functions connect accountable ownership, understanding of the clinical context, evidence about system behavior, and decisions about risk treatment. The proposed application is an interpretation of the framework, rather than a HIPAA compliance certification. (NIST AI Risk Management Framework 1.0.)
Healthcare-specific transparency expectations reinforce the need for documentation. Within its defined certification scope, ONC’s decision support intervention criterion addresses source information for predictive interventions, including intended use, training data relevance, performance, and risk management. It is not a universal approval requirement for every healthcare AI application. Nevertheless, these information categories provide useful questions for organizations assessing a system. (ONC decision support intervention criterion.)
Implementation can start with a bounded workflow and a small number of well-understood sources. Map the decision, assign owners, establish the permitted data uses, and measure the quality of those sources. Then implement context-preserving transformations, enforce access, and test the application with representative cases and failure scenarios. Establish release criteria before the pilot produces pressure to expand.
Expansion should depend on demonstrated benefit and an ability to maintain the foundation. Monitor source latency, missingness, identity conflicts, mapping changes, access failures, retrieval quality, and application outcomes. Material changes to a source system, terminology, patient population, model, or workflow should trigger reassessment. The organization also needs authority to pause an unsafe application and a usable clinical fallback.
A clinician reviewing an AI-generated account should be able to determine which patient and encounter it concerns, what evidence supports it, how current that evidence is, and where information remains incomplete or contradictory. The organization should be able to explain who authorized the processing, which services handled the information, and what happens when an error is discovered. Building those capabilities into the data foundation gives healthcare AI a basis for earned, continually tested trust.