AI Compliance

HIPAA-Compliant AI Development Checklist: When a BAA Isn't Enough

A HIPAA compliant AI development checklist built from HHS rules: BAA scope, PHI data flows, de-identification, audit logs and the gaps teams miss.

Cover image for an article about a HIPAA-compliant AI development checklist covering business associate agreements, PHI data flows and audit logs
In this article
  1. What "HIPAA-compliant AI" actually means
  2. Step 1: Decide what data the feature actually needs
  3. Step 2: Get the right BAA, then read what it excludes
  4. Step 3: Map every hop the PHI takes
  5. Step 4: Apply minimum necessary to prompts and retrieval
  6. Step 5: Map Security Rule safeguards onto the AI stack
  7. Step 6: Log enough to reconstruct any AI decision
  8. Step 7: Put the AI system in your risk analysis and write it down
  9. Step 8: Control training, fine-tuning and secondary use
  10. Step 9: The rules that are not HIPAA
  11. The HIPAA compliant AI development checklist
  12. Five ways teams fail this after signing the BAA
  13. What to do in the next two weeks
  14. Frequently asked questions
  15. Sources

HIPAA does not mention artificial intelligence anywhere in its text, and no agency certifies AI tools as compliant. Microsoft says it directly in its own compliance documentation: "There's currently no certification standard that the Department of Health and Human Services approves to demonstrate compliance with HIPAA or the HITECH Act by a business associate." So a useful HIPAA compliant AI development checklist is not a list of approved vendors. It is a list of decisions you have to make, implement and write down about which data the system touches, who is contractually on the hook, what controls sit around it, and what you can prove after the fact.

The short answer: get the right business associate agreement with every vendor in the path, decide deliberately whether the feature needs PHI at all, apply minimum necessary to prompts and retrieval rather than just to database tables, log every AI decision well enough to reconstruct it, put the AI system inside your Security Rule risk analysis, and keep the documentation for six years.

This article walks through ten steps in that order, with the regulatory citation behind each one, a comparison of what the major AI platform BAAs actually cover as of September 2026, and the specific places where teams that did sign a BAA still ended up out of compliance. It is a practitioner's guide, not legal advice; confirm any scope decision with your privacy officer and counsel.

What "HIPAA-compliant AI" actually means

Three facts frame everything else.

First, HIPAA regulates uses and disclosures of protected health information, not technologies. Nothing in the Privacy Rule or Security Rule turns on whether a system uses a large language model. An AI feature that touches PHI is in scope exactly as a reporting query is in scope.

Second, any vendor that creates, receives, maintains or transmits PHI on your behalf is a business associate under 45 CFR 160.103, and so is any subcontractor that does the same on the business associate's behalf. A model provider, a vector database host, a tracing platform and an error-reporting service are all business associates if PHI reaches them. The definition does not care that the PHI arrived inside a prompt.

Third, HHS has already said out loud that AI-specific data is covered. In the January 2025 proposal to modernize the Security Rule (90 FR 898), the Department wrote that "ePHI, including ePHI in AI training data, prediction models, and algorithm data that is maintained by a regulated entity for covered functions is protected by the HIPAA Rules and all applicable standards and specifications." It went further on process: "we expect that a regulated entity interested in using AI would include the use of such tools in its risk analyses and associated risk management activities," and the risk analysis "must include consideration of, among other things, the type and amount of ePHI accessed by the AI tool, to whom the data is disclosed, and to whom the output is provided."

Key takeaway: There is no AI exception and no AI-specific rule. The obligations you already have apply, and HHS has stated in writing that training data, model artifacts and algorithm data are inside the perimeter.

That proposal is still a proposal. No final rule has been published, and the Unified Agenda now lists the rulemaking under Long-Term Actions with a projected final action of July 2027. The current Security Rule text is what binds you today; the proposal is the clearest statement of how HHS reads that text when AI is involved.

Step 1: Decide what data the feature actually needs

Most compliance cost in healthcare AI is created in the first design meeting, by defaulting to full records because they were easy to query.

The Privacy Rule gives you four tiers to choose from, and they carry very different burdens.

Decision tree with three questions leading to PHI, a limited data set, Safe Harbor de-identification or expert determination
Which data the feature needs, decided before any code is written. Methods and identifier lists from 45 CFR 164.514.

Full PHI. Required when the feature has to identify or act on a specific patient: a scheduling agent, a prior-authorization assistant, a chart summarizer. Everything in this article applies.

A limited data set. Under 45 CFR 164.514(e) you strip 16 named direct identifiers but keep dates, city, state, ZIP code and ages. It is still PHI, still needs safeguards, and may only be used for research, public health or health care operations, under a data use agreement in which the recipient agrees not to re-identify or contact individuals. This is often the right tier for model evaluation and analytics.

Safe Harbor de-identification. Remove all 18 identifier categories listed at 164.514(b)(2), and hold no actual knowledge that what remains could identify someone. Done correctly, the result is not PHI and HIPAA does not apply to it.

Expert determination. A person with appropriate statistical expertise determines and documents that the risk of identification is very small. HHS guidance is explicit that there is "no explicit numerical level of identification risk that is deemed to universally meet the 'very small' level," and notes that practitioners often issue time-limited certifications because re-identification capability improves over time.

The trap is free text. HHS de-identification guidance states that "the de-identification standard makes no distinction between data entered into standardized fields and information entered as free text (i.e., structured and unstructured text) -- an identifier listed in the Safe Harbor standard must be removed regardless of its location in a record if it is recognizable as an identifier." It adds that clinical narratives "are information rich and may provide context that readily allows for patient identification." A regex that catches names in a structured field and misses them in a discharge summary has not de-identified anything.

Practical rule: run your de-identification pipeline over a few hundred real notes and have a human review the output before you rely on it. Measure the miss rate. If you cannot get it acceptably low, treat the data as PHI and build the controls, which is the honest and cheaper outcome than discovering the gap later.

Step 2: Get the right BAA, then read what it excludes

Every AI vendor in the path needs a business associate agreement executed before the first record moves. 45 CFR 164.504(e)(2) lists what the contract must require, including that the business associate will not use or disclose PHI other than as permitted, will use appropriate safeguards and comply with the Security Rule, will report breaches, will bind its own subcontractors to the same terms, will support individual access and amendment rights and accounting of disclosures, will make its books available to the Secretary, and will return or destroy PHI at termination if feasible.

Two things about that list matter for AI specifically. The subcontractor flow-down means your model provider's own hosting and monitoring vendors are inside the chain. And the return-or-destroy clause is hard to satisfy for anything a model has already been fine-tuned on, which is a reason to keep training data separate and deletable.

Getting a signature is the easy part. The hard part is that AI vendor BAAs are scoped to particular products and, increasingly, to particular features.

Platform BAA available (Sept 2026) Scope you have to check
AWS Yes; AWS states you agree not to use HIPAA Eligible Services with PHI "without first entering into an AWS business associate agreement" Only HIPAA Eligible Services count. Amazon Bedrock is listed; Amazon SageMaker is listed but excludes Studio Lab, Ground Truth Plus, and Public and Vendor Workforce. AWS adds that "Customers still must configure these services consistent with HIPAA requirements"
Microsoft Azure Yes; the Microsoft HIPAA BAA is available through the Online Services Data Protection Addendum by default to covered entity and business associate customers, and Azure is an in-scope service Flagged prompts and completions can be stored for human review by authorized Microsoft employees unless you are approved for modified abuse monitoring, which requires meeting Limited Access criteria and applying. Preview features may follow different practices
OpenAI Yes, on named HIPAA-eligible products: ChatGPT for Healthcare, ChatGPT for Enterprise with Regulated Workspace, ChatGPT FedRAMP, ChatGPT for Clinicians, API with Modified Retention, API FedRAMP with Modified Retention Coverage is feature-level. OpenAI publishes a list of functionality not covered by the BAA and states that "Features not listed here are not automatically covered under the BAA"
Anthropic Yes, for HIPAA-ready services such as the first-party API or Enterprise plans; the organization's Primary Owner activates HIPAA compliance or signs and requests it The BAA "only covers the single organization that accepted it, and excludes features such as Claude Console, Claude Cowork, or features currently in beta." Covered Models require 30-day data retention and are not available with zero data retention. Consumer plans are out of scope
Google Cloud Yes; customers subject to HIPAA who want to use any Google Cloud product with PHI "must review and accept Google's Business Associate Agreement" The BAA covers Google Cloud infrastructure plus the listed in-scope services, and "is not subject to modification." Google states the covered entity "is responsible for building a HIPAA compliant solution using the approved Google Cloud services"

Read that table as one message: the BAA moves the boundary, it does not remove the work. Microsoft answers the question directly in its own FAQ. Asked whether having a BAA with Microsoft ensures compliance, the answer is "No. By offering a Business Associate Agreement, Microsoft helps support your HIPAA compliance. However, using Microsoft services doesn't on its own achieve HIPAA compliance."

One more operational detail: hyperscalers will not sign your BAA template. Microsoft states it cannot use a customer's BAA because its services are standardized; Google states its BAA is not subject to modification. Budget review time for their paper, not negotiation time for yours.

Step 3: Map every hop the PHI takes

OCR's own remediation advice, published with its April 23, 2026 settlements, starts with mapping: "Identify where ePHI is located in the organization, including how ePHI enters, flows through, and leaves the organization's information systems."

In an AI system that map is longer than teams expect, because the interesting parts are the side paths rather than the request path.

Diagram of a four-step AI request pipeline plus six side paths that also carry protected health information
The four hops teams scope, and the six they forget. Each side path is a separate vendor, BAA and retention decision.

Walk the list and answer three questions for each box: does PHI reach it, is there a BAA, and how long does it keep the data?

  • The vector store. Chunked notes are PHI. So is the filter metadata. Embeddings are derived from PHI and are not a de-identification method.
  • Your own prompt and completion logs. Debug logging that captures the full prompt creates a new ePHI store, usually with wider access than the source system.
  • Provider-side retention. Defaults differ by product and tier, and the setting you need may be an application rather than a toggle. Azure's modified abuse monitoring requires approval. OpenAI's HIPAA-eligible API path is specifically "API with Modified Retention." Anthropic's Covered Models require 30-day retention and cannot be run with zero data retention.
  • LLM observability and tracing. Capturing full prompts is the product's purpose. Treat it as a business associate or configure redaction before the span leaves your process.
  • Error and crash reporting. Stack traces and request bodies carry the payload that broke. This is the leak that survives every review because nobody thinks of it as a data store.
  • Evaluation and test datasets. Real cases copied into a spreadsheet or a laptop to build a golden set is the most common quiet disclosure in healthcare AI projects.
  • Product analytics and session replay. Session recording tools capture what the clinician typed.

Write this down as a table with owner, vendor, BAA date, retention period and region. That table is the artifact an auditor asks for, and it is the input to your risk analysis.

Step 4: Apply minimum necessary to prompts and retrieval

The minimum necessary standard at 45 CFR 164.514(d) requires you to identify who needs access to what, limit it accordingly, and limit each request for PHI to what is reasonably necessary. Paragraph (d)(5) is blunt about the failure mode: you may not use, disclose or request an entire medical record "except when the entire medical record is specifically justified as the amount that is reasonably necessary."

In a retrieval-augmented system this becomes three concrete engineering requirements.

Filter retrieval by the user's own permissions, not just by relevance. The retriever must apply the same access rules as the source system. A semantic search that can surface any patient's note to any logged-in user is an impermissible disclosure waiting for a query, and it will look exactly like a feature in your logs.

Trim the prompt. Send the fields the task needs. If the task is "draft a referral letter," it does not need the full problem list, the billing history and ten years of encounters. Trimming reduces both compliance surface and token cost, which makes it an easy sell internally. Our write-up on RAG chatbot development cost covers the retrieval side of this in more detail.

Constrain the output. Decide what the model is allowed to include, and check it. A summary that helpfully restates the patient's address to a user not entitled to see it is still a disclosure.

Step 5: Map Security Rule safeguards onto the AI stack

The Security Rule is technology-neutral, which makes it easy to nod at and hard to apply. Here is the translation for an AI system, with the citation for each control.

Security Rule requirement Citation What it means for an AI system
Risk analysis (Required) 164.308(a)(1)(ii)(A) The AI system, its vendors, its prompt logs and its training data are all in the assessment
Risk management (Required) 164.308(a)(1)(ii)(B) Written decisions on the risks you found, including accepted ones
Information system activity review (Required) 164.308(a)(1)(ii)(D) Someone actually reads the AI access logs on a schedule
Sanction policy (Required) 164.308(a)(1)(ii)(C) Pasting PHI into an unapproved AI tool is a named violation with a named consequence
Access control 164.312(a)(1) Access limited to persons "or software programs" granted rights. Service accounts and agent identities are in scope
Unique user identification (Required) 164.312(a)(2)(i) Every AI call is attributable to a human or a named service identity, never a shared key
Audit controls 164.312(b) Mechanisms that record and examine activity in systems containing ePHI
Integrity 164.312(c)(1) Protect ePHI from improper alteration. An agent that can write to a chart needs change control and reversibility
Person or entity authentication 164.312(d) Verify who or what is calling, including machine-to-machine calls from agents
Encryption in transit and at rest (Addressable) 164.312(a)(2)(iv), (e)(2)(ii) Implement it, or document a reasonable alternative and why. There is rarely a defensible alternative in 2026
Documentation retention (Required) 164.316(b)(2)(i) Keep the policies, the risk analysis and the decisions for six years

Two nuances repay attention.

"Addressable" does not mean optional. It means implement the specification, or document why it is not reasonable and appropriate and what you did instead. HHS proposed in 2025 to "remove the distinction between 'addressable' and 'required' implementation specifications," explaining that the aim is to make clear "what has always been a requirement." Treat addressable specifications as requirements with a documentation escape hatch you will probably never want to use.

And access control explicitly covers "software programs," not just people. That single phrase is why agent identities, tool credentials and service accounts belong in your access review. If you are building agents that call internal systems, the controls in MCP server development work are the same controls the Security Rule is asking about.

For a control-by-control walkthrough, NIST SP 800-66 Revision 2 (February 2024) is the practical companion to the rule text. It is guidance, not law, but OCR investigators and NIST speak the same vocabulary.

Step 6: Log enough to reconstruct any AI decision

Audit controls at 164.312(b) require mechanisms that record and examine activity in systems holding ePHI, and 164.308(a)(1)(ii)(D) requires you to actually review those records. For an AI feature, "activity" has to include the model's behavior, not just the database read.

A workable minimum for each AI interaction:

  1. Timestamp, and the identity of the human user and the service account.
  2. The patient records or record IDs retrieved, and the access rule that allowed each one.
  3. A hash or reference to the exact prompt sent, plus the prompt template version.
  4. The model name and version, and the inference parameters.
  5. The output returned, or a reference to it.
  6. Any tool or write action the system took, with its result.
  7. The human decision that followed: accepted, edited or rejected.

Item 7 is the one teams skip and regulators and litigators care about most, because it is the only evidence that human oversight was real rather than nominal.

Key takeaway: If your audit log contains prompts or outputs, that log is now an ePHI system. It needs its own access controls, encryption, retention limit and BAA coverage. Many teams solve this by storing hashes and references in the wide-access log and keeping the content in the same store as the source records.

Retention deserves a deliberate decision. The Security Rule's six-year clock at 164.316(b)(2)(i) applies to documentation of required actions and assessments, not to every application log. Keeping raw prompts forever "just in case" enlarges your breach surface with no compliance benefit.

Step 7: Put the AI system in your risk analysis and write it down

This is the step OCR is enforcing hardest, and it has nothing to do with AI specifically.

On April 23, 2026, OCR announced settlements with four regulated entities following ransomware investigations, totaling $1,165,000 and affecting more than 427,000 individuals, with corrective action plans monitored for two years. All four involved a failure to conduct an accurate and thorough risk analysis. OCR described the announcement as marking "19 completed investigations from ransomware breaches and 13 completed investigations in OCR's Risk Analysis Initiative." OCR Director Paula M. Stannard framed the lesson: "Proactively implementing the HIPAA Security Rule before a breach or an OCR investigation not only is the law but also is a regulated entity's best opportunity to prevent or mitigate the harmful effects of a successful cyberattack."

Note who was in that group: a provider network, an imaging company, a third-party administrator acting as a business associate, and a self-funded employer health plan. Enforcement is not limited to hospitals.

For an AI feature, the risk analysis addendum HHS described in the 2025 proposal is short and specific. Document the type and amount of ePHI the tool accesses, who it is disclosed to, and who receives the output. Then add the AI-specific failure modes that a generic template will miss:

  • The model returns a confident wrong answer and a clinician or coder acts on it.
  • Prompt injection in an ingested document redirects an agent to retrieve or send data it should not.
  • Retrieval leaks another patient's record because a filter was applied in the application layer but not in the index.
  • A provider-side retention default was never changed, so prompts persist outside your control.
  • Fine-tuning data is memorized and surfaced in an output. The HHS proposal cites exactly this risk, noting that "generative AI tools have produced in their output the names and personal information of persons included in the tools' sources of training data."

The stakes for the documentation itself are set by 45 CFR 102.3. For violations due to willful neglect that are not corrected, the inflation-adjusted minimum penalty is $73,011 per violation with an annual cap of $2,190,294 per identical violation (2025 adjusted amounts, as amended in January 2026). The willful-neglect tiers exist to punish the absence of a program, which is why an imperfect documented analysis is worth far more than a perfect undocumented one.

Step 8: Control training, fine-tuning and secondary use

Three distinct questions hide inside "can we train on our data?"

Will the vendor train on your inputs? The major enterprise AI platforms say no by default for business tiers, and their HIPAA-eligible configurations tighten it further. Get the specific product's current statement in writing and attach it to your vendor file, because product tiers change.

Will you fine-tune on PHI? That is a use of PHI requiring a permitted purpose, minimum necessary limits, access controls and a deletion path. If the fine-tuned weights cannot be untangled from the training data, your BAA's return-or-destroy obligation becomes awkward at termination. Prefer retrieval over live records, or fine-tune on a de-identified or limited data set.

What happens to evaluation data? Golden datasets are the least governed and most copied artifact in any AI project. Put them in the same store as the source records, or de-identify them properly and document the method. "It's just test data" has never survived a breach investigation.

Our comparison of private LLM options versus ChatGPT Enterprise walks through the deployment choices behind these questions, including when self-hosting genuinely reduces compliance surface and when it just moves the work.

Step 9: The rules that are not HIPAA

HIPAA is a floor for privacy and security. It says nothing about whether the AI is any good, and three other regimes do.

FDA device regulation. Software that supports clinical decisions can be a medical device. The statutory carve-out at 21 U.S.C. 360j(o)(1)(E) excludes software intended for displaying medical information, supporting or providing recommendations to a health care professional about prevention, diagnosis or treatment, and "enabling such health care professional to independently review the basis for such recommendations that such software presents so that it is not the intent that such health care professional rely primarily on any of such recommendations." It also does not apply at all if the function acquires, processes or analyzes a medical image or a signal from an in vitro diagnostic device. That independent-review clause is where LLM features most often fall out of the exclusion: if the clinician cannot see and check the basis, or is expected to rely on the output, you may be building a device. FDA issued final Clinical Decision Support Software guidance in January 2026 (docket FDA-2017-D-6569), and it is worth reading before you ship anything that makes a recommendation.

Section 1557 nondiscrimination. 45 CFR 92.210 prohibits covered entities from discriminating through the use of patient care decision support tools, and imposes "an ongoing duty to make reasonable efforts to identify" tools that use input variables measuring race, color, national origin, sex, age or disability, plus a duty to make reasonable efforts to mitigate the resulting risk. The 2024 final rule set compliance with paragraphs (b) and (c) at within 300 days of the July 5, 2024 effective date, which is May 1, 2025. A 2025 court order vacated other parts of that rule only to the extent they expanded Title IX's definition of sex discrimination to include gender-identity discrimination; HHS confirmed in June 2026 that "the other provisions of the Section 1557 Rule remain in force." In practice, 92.210 means a written inventory of decision support tools, a note on which sensitive variables each uses, and evidence you looked for disparities.

State AI disclosure laws. California requires a health facility, clinic, physician's office or group practice using generative AI to produce written or verbal patient communications about clinical information to include a disclaimer that the communication was AI-generated, plus clear instructions for reaching a human. The requirement does not apply where a licensed or certified provider reads and reviews the communication. Texas requires that where an AI system is used in relation to health care service or treatment, the provider disclose that to the recipient no later than the date the service is first provided, except in emergencies; that statute took effect January 1, 2026. Both are narrow, and both are easy to fail by shipping a patient-facing assistant without a disclaimer.

The HIPAA compliant AI development checklist

One page. Every line is something you can mark done and point at evidence for.

Before you build

  1. Named owner for the AI system, and a named privacy or security reviewer.
  2. Written statement of purpose: what decision it influences, who sees the output.
  3. Data tier decision recorded: PHI, limited data set, Safe Harbor, or expert determination.
  4. If de-identifying, method documented and miss rate measured on real records including free text.
  5. Data flow table listing every vendor, with owner, BAA status, retention period and data region.
  6. BAA executed with every vendor before the first record moves, including subcontractor flow-down.
  7. Product and feature scope of each BAA confirmed against the vendor's published coverage list.
  8. Provider-side retention and human review settings configured or formally applied for.

In the build

  1. Retrieval enforces the user's own access rights inside the index, not only in the application.
  2. Prompts carry only the fields the task requires; no whole-record dumps.
  3. Output constraints defined and tested for what the model may include.
  4. Encryption in transit and at rest, with key management documented.
  5. Unique identity for every human and service caller; no shared API keys.
  6. Audit log capturing the seven fields in step 6, including the human decision.
  7. Audit log itself treated as an ePHI store, with its own access controls and retention.
  8. Prompts, traces and stack traces excluded from tools that lack a BAA, or redacted before egress.
  9. Evaluation and test datasets held in a governed store, never on laptops or in shared spreadsheets.

Before you go live

  1. Risk analysis updated to include the AI system, its vendors and its data stores.
  2. AI-specific failure modes assessed: wrong outputs, prompt injection, cross-patient retrieval leakage, memorization.
  3. Accuracy and safety evaluation run on a held-out set, with results recorded and a pass threshold agreed.
  4. Human review step designed where the output affects care, coding or coverage, and logged when used.
  5. Workforce training covering this system specifically, plus a sanction policy naming unapproved AI tools.
  6. FDA device analysis documented if the tool makes recommendations to clinicians.
  7. Section 1557 decision support tool inventory entry and sensitive-variable note, if you are a covered entity.
  8. State AI disclosure requirements checked for every state you operate in.
  9. Incident runbook covering an AI-specific breach, with the 60-day clocks written into it.

Ongoing

  1. Scheduled review of AI access logs, with the review itself recorded.
  2. Re-run the risk analysis when the model, prompt, data source or vendor changes.
  3. Annual confirmation that each vendor's BAA scope still covers the features you use.
  4. All of the above retained for six years per 164.316(b)(2)(i).

Five ways teams fail this after signing the BAA

Feature drift inside a covered product. The workspace has a BAA. Then someone enables a connector, a browsing tool or a beta feature that the BAA's published coverage list excludes. The contract did not change; the data flow did. This is why item 29 exists.

The observability tool nobody put in the inventory. Tracing was added during a debugging week, captures full prompts, and was never reviewed. Same story for error reporting and session replay.

Retrieval permissions enforced in the wrong layer. The UI filters correctly, the index does not, and an unusual query returns another patient's note. The application looks like it is working.

The evaluation set. A clinician exports 300 real cases to a spreadsheet to grade model answers. The intent is quality; the act is a disclosure to an ungoverned store.

A risk analysis that predates the AI system. The document exists, was thorough in its day, and does not mention the tool that now reads charts. Under 164.316(b)(2)(iii) documentation must be reviewed and updated in response to operational changes. A new AI system is an operational change.

What to do in the next two weeks

If you are starting from zero, the order matters more than the speed.

Week one: inventory. List every place an AI feature already touches PHI, including the ones that arrived without a review. Add the side paths from the diagram above. For each, record the vendor, whether a BAA exists, and the retention period. You will find at least one surprise, and finding it now is the entire point.

Week two: close the two widest gaps. In most organizations those are provider-side retention settings that were never configured and logging tools that were never scoped. Both are configuration changes rather than projects.

Then update the risk analysis and write down what you decided, including the risks you are accepting and why. The documentation is not bureaucratic overhead; under the enforcement pattern OCR has established, it is the difference between a corrective action plan and a willful neglect finding.

Fleurant AI builds AI for regulated industries, including the PHI data-flow mapping, audit logging, evaluation harnesses and control documentation that make a healthcare AI system defensible, alongside custom AI and AI agent work. If you want a second opinion on a specific architecture before it goes live, talk to a specialist and we will walk the data flow with you. For anything that turns on legal interpretation, your counsel and compliance team make the call.

Frequently asked questions

Is ChatGPT HIPAA compliant?

No product is "HIPAA compliant" by itself, and consumer ChatGPT is not covered by a BAA at all. OpenAI does offer a BAA for specific products, including ChatGPT for Healthcare, ChatGPT for Enterprise with a Regulated Workspace, ChatGPT for Clinicians and the API with Modified Retention. Even inside those products, OpenAI publishes a list of features that the BAA does not cover and states that features not listed as covered are not automatically covered.

Is a signed BAA enough to make an AI tool HIPAA compliant?

No. A BAA allocates responsibility; it does not implement controls. Microsoft says so plainly in its own HIPAA documentation: having a BAA with Microsoft does not ensure your compliance, and your organization remains responsible for its compliance program and for how it uses the service. You still owe a risk analysis, access controls, audit logs, minimum necessary limits and documentation.

Can we use patient data to train or fine-tune a model?

Sometimes, but it is a use of PHI that needs a legal basis and controls, not a technical detail. HHS stated in its 2025 Security Rule proposal that ePHI in AI training data, prediction models and algorithm data is protected by the HIPAA Rules. Treat fine-tuning sets as a PHI system of record: minimum necessary, access controls, retention limits and a documented purpose. Many teams find de-identified data or retrieval over live records is enough.

Are BAAs from no-code and automation platforms enough for PHI?

A BAA is necessary but rarely sufficient with these platforms, because PHI usually passes through connectors, logs and execution histories you cannot fully inspect or purge. Ask where run logs live, how long they persist, whether support staff can read them, and which subprocessors are involved. If the vendor cannot answer in writing, you cannot document the data flow, and you cannot pass a risk analysis.

Does de-identified data take us outside HIPAA?

Yes, if it is genuinely de-identified under 45 CFR 164.514(b) by Safe Harbor or expert determination. The hard part is free text. HHS guidance states that the standard makes no distinction between structured fields and free text, so an identifier must be removed wherever it appears in a record. Clinical narratives are also rich enough in context that they can allow identification even after names are stripped.

Does HIPAA require us to encrypt PHI in an AI system?

Encryption is currently an "addressable" implementation specification at 45 CFR 164.312(a)(2)(iv) and (e)(2)(ii), which means you must implement it or document a reasonable alternative and why. In practice, for a modern cloud AI system there is no defensible alternative, and HHS has proposed removing the addressable and required distinction entirely. Encrypt in transit and at rest, and write down your decision either way.

Do we have to tell patients an AI was involved?

HIPAA does not require it, but other rules may. California requires a disclaimer and instructions for reaching a human when generative AI produces patient communications about clinical information, unless a licensed provider reads and reviews the message. Texas requires providers to disclose AI use in relation to health care service or treatment no later than the date the service is first provided. Check your states.

Who is liable if the AI vendor causes the breach?

Both parties have obligations. A business associate must notify the covered entity of a breach of unsecured PHI without unreasonable delay and no later than 60 calendar days after discovery, and the covered entity then owes notice to individuals on its own 60-day clock. A covered entity that knew of a pattern of the business associate breaching the BAA and did not act is also out of compliance. Contracts do not transfer your exposure.

Sources

  1. 45 CFR 160.103: Definitions (including "business associate"), Electronic Code of Federal Regulations
  2. 45 CFR 164.504: Uses and disclosures: Organizational requirements (business associate contracts), Electronic Code of Federal Regulations
  3. 45 CFR 164.514: Other requirements relating to uses and disclosures of protected health information, Electronic Code of Federal Regulations
  4. 45 CFR 164.308: Administrative safeguards, Electronic Code of Federal Regulations
  5. 45 CFR 164.312: Technical safeguards, Electronic Code of Federal Regulations
  6. 45 CFR 164.316: Policies and procedures and documentation requirements, Electronic Code of Federal Regulations
  7. 45 CFR 164.402: Breach Notification Rule definitions, Electronic Code of Federal Regulations
  8. 45 CFR 164.410: Notification by a business associate, Electronic Code of Federal Regulations
  9. 45 CFR 102.3: Penalty adjustment and table, Electronic Code of Federal Regulations
  10. 45 CFR 92.210: Nondiscrimination in the use of patient care decision support tools, Electronic Code of Federal Regulations
  11. HIPAA Security Rule To Strengthen the Cybersecurity of Electronic Protected Health Information (90 FR 898), Federal Register
  12. Unified Agenda entry for RIN 0945-AA22, HIPAA Security Rule, Office of Information and Regulatory Affairs, reginfo.gov
  13. HHS' Office for Civil Rights Settles Four HIPAA Security Rule Ransomware Investigations, U.S. Department of Health and Human Services
  14. Guidance Regarding Methods for De-identification of Protected Health Information, U.S. Department of Health and Human Services, Office for Civil Rights
  15. HIPAA Eligible Services Reference, Amazon Web Services
  16. Health Insurance Portability and Accountability Act (HIPAA) & HITECH Act, Microsoft Learn
  17. Data, privacy, and security for Foundry Models sold by Azure, Microsoft Learn
  18. Foundry Models sold by Azure abuse monitoring, Microsoft Learn
  19. HIPAA eligible products and functionality, OpenAI Help Center
  20. Business Associate Agreements (BAA) for Commercial Customers, Anthropic Privacy Center
  21. HIPAA compliance offering, Google Cloud
  22. Implementing the HIPAA Security Rule: A Cybersecurity Resource Guide (NIST SP 800-66r2), National Institute of Standards and Technology
  23. Clinical Decision Support Software: Guidance for Industry and FDA Staff (January 2026), U.S. Food and Drug Administration
  24. 21 U.S.C. 360j(o): Regulation of medical and certain decision support software, Office of the Law Revision Counsel, U.S. House of Representatives
  25. Assembly Bill 3030, Health care services: artificial intelligence (Chapter 848, 2024), California Legislative Information
  26. Texas House Bill 149, Texas Responsible Artificial Intelligence Governance Act (enrolled), Texas Legislature Online
  27. Notice of Vacatur Regarding Certain Provisions of the 2024 Nondiscrimination in Health Programs and Activities Final Rule (91 FR 32887), Federal Register

Free, no-obligation consultation

Have a question about HIPAA-compliant AI development?

Tell us what you're working on or what you'd like to know. A specialist will get back to you with practical next steps, whether or not we end up working together.

  1. 1Send your question or project details (takes 2 minutes)
  2. 2A specialist reviews it and replies within 1 business day
  3. 3Get clear, practical next steps, free