AI Compliance

AI Audit Trail Requirements: What to Log for Every AI Decision

AI audit trail requirements as of October 2026: which rules apply, a field-by-field logging spec for LLM calls and agent actions, retention periods and a test.

Cover image for an article explaining AI audit trail requirements and what to log for every LLM call and AI agent decision
In this article
  1. What AI audit trail requirements actually ask for
  2. Which US rules require an AI audit trail?
  3. What to log for every LLM call: a field-by-field specification
  4. What to log for AI agent decisions
  5. Industry add-ons: banking, insurance, healthcare and broker-dealers
  6. How long to keep AI audit logs
  7. How to make an AI audit trail tamper-evident
  8. What not to put in your AI audit logs
  9. The AI audit trail reconstruction test
  10. Common AI audit trail mistakes and how to avoid them
  11. Conclusion: build the decision record first
  12. Frequently asked questions
  13. Sources

AI audit trail requirements come down to one test: months from now, can you show exactly who asked an AI system to do something, what it saw, what it produced, what action followed, who approved it, and prove that record has not been changed? No single US law spells out an AI log format. Instead, the obligation arrives through rules you already follow, including HIPAA's audit controls, Regulation B's adverse action rules, broker-dealer recordkeeping and the NAIC model bulletin for insurers, plus a new wave of state laws on automated decisions that start applying in 2027.

This guide gives you three things you can use this quarter. First, a map of which rules reach AI decision records, with retention periods, current as of October 2026. Second, a field-by-field logging specification for LLM calls and for AI agents that take actions. Third, a reconstruction test you can run against any AI system to find out whether its audit trail would survive an examiner, a plaintiff's lawyer or an angry customer.

It is written for the founders, CTOs, compliance officers and operations leaders who will be asked to produce the evidence. It is not legal advice: confirm the rules that apply to your systems with your counsel or compliance team.

What AI audit trail requirements actually ask for

Strip away the vendor language and every regulator is asking the same three questions about an AI-influenced decision.

  1. What happened? The inputs, the model, the output and any action taken, in order.
  2. Why did it happen? The data, documents, rules and human judgments that shaped the result, in a form you can explain to the affected person.
  3. Who is accountable? The named human or role who requested, reviewed, approved or overrode it, and the owner of the system.

A good baseline for "what happened" already exists. NIST SP 800-53 control AU-3, Content of Audit Records, says every audit record should establish what type of event occurred, when, where, its source, its outcome and the identity of the people or entities involved. That list was written for ordinary IT systems. AI systems need the same six elements plus things a database log never had to capture: the prompt, the retrieved context, the exact model version and the reasoning the system offered.

The difference between a log and an audit trail is purpose. Application logs exist so engineers can fix problems this week. An audit trail exists so someone else can reconstruct a decision years later, without trusting you. That changes what you capture, how long you keep it and who can delete it.

Key takeaway: If you cannot rebuild a single AI decision end to end from your records, with the identity of every human involved, you do not have an audit trail. You have telemetry.

Which US rules require an AI audit trail?

The table below maps the rules most likely to reach a US small or mid-sized company using AI in a regulated workflow. None of them mentions "prompts." All of them reach the records an AI system creates or influences.

Rule Who it reaches What it asks for that touches AI records Retention floor
HIPAA Security Rule, 45 CFR 164.312(b) and 164.316 Covered entities and business associates handling ePHI Mechanisms that record and examine activity in systems that contain or use ePHI; documented policies and actions 6 years for required documentation
Regulation B (ECOA), 12 CFR 1002.9 and 1002.12 Creditors, including lenders using AI in underwriting Specific principal reasons for adverse action; records of the application and notices 25 months (consumer credit), 12 months (most business credit)
SEC Rule 17a-4(b)(4) plus FINRA supervision rules Broker-dealers Originals of communications received and copies of communications sent relating to the business 3 years, first 2 easily accessible
NAIC AI model bulletin (adopted by many states) Insurers using AI in regulated practices An AIS program covering data lineage, traceability, auditability and "data and record retention" Set by the state and the insurer's program
Colorado SB 26-189 Developers and deployers of automated decision-making technology used in consequential decisions Records needed to demonstrate compliance; post-decision explanations of the technology's role At least 3 years
California CCPA ADMT regulations Businesses using ADMT for significant decisions about California consumers Pre-use notice, access and opt-out rights; risk assessments Compliance from January 1, 2027
EU AI Act, Articles 12, 19 and 26(6) Providers and deployers of high-risk AI systems in the EU market Automatic event logging over the system's lifetime At least 6 months

A few of these deserve more detail, because the commonly repeated versions online are out of date.

Colorado replaced its AI Act before it took effect

Colorado's original law, SB 24-205, was delayed to June 30, 2026, and then replaced. The legislature passed SB 26-189, signed May 14, 2026, which repeals and reenacts the 2024 provisions around "automated decision-making technology" used in consequential decisions about education, employment, housing, lending, insurance, health care and essential government services. According to the General Assembly's summary, developers and deployers must keep records needed to demonstrate compliance for at least three years, deployers must give a plain-language description of the technology's role within 30 days after an adverse consequential decision, and consumers get rights to access the personal data used, correct factual errors and obtain meaningful human review. Developer documentation duties begin January 1, 2027, and the attorney general must adopt rules on the post-decision disclosures by the same date. Check the enrolled act for the start date of each deployer duty.

Notice what that 30-day explanation requires operationally. You cannot describe the role an AI system played in a specific decision unless you recorded, at the time, which system ran, what it received and what it returned. That is an audit trail requirement, even though the summary never uses the phrase.

California's ADMT rules start in January 2027

The California Privacy Protection Agency finalized its regulations on automated decisionmaking technology, risk assessments and cybersecurity audits in September 2025. The regulations took effect January 1, 2026, and businesses that use ADMT to make significant decisions must comply with the ADMT requirements beginning January 1, 2027. Risk assessment attestations and summaries are due to the agency by April 1, 2028. Access requests about ADMT are hard to answer without decision-level records.

Lending: the reasons must be specific

Regulation B requires that a statement of reasons for adverse action be specific and indicate the principal reasons. The text of 12 CFR 1002.9(b)(2) says it is insufficient to state that the decision was based on internal standards or that the applicant failed to achieve a qualifying score. If an AI model or agent contributes to a credit decision, your log has to retain the factors behind its output, not just the score. Section 1002.12 then sets the retention period at 25 months for consumer credit.

Banking: model risk guidance now excludes generative AI, but not the duty

The federal banking agencies' revised model risk guidance, SR 26-2, places generative and agentic AI outside its scope and hands those decisions back to each bank's own risk management. We covered the details in what changed for bank AI under SR 26-2. The practical effect is that your own policy, not a supervisory template, now defines what you log for those systems. That makes a written logging standard more important, not less.

Broker-dealers: agents are now named as a risk

FINRA's 2026 Annual Regulatory Oversight Report discusses AI agents for the first time. It flags agents acting without human validation, acting beyond the user's intended scope and authority, and multi-step reasoning that makes outcomes hard to trace or explain. Among the practices it lists for firms to consider are storing prompt and output logs, tracking which model version was used and when, tracking agent actions and decisions, and defining where human-in-the-loop oversight applies. Separately, SEC Rule 17a-4(b)(4) requires broker-dealers to keep communications relating to the business for three years. A message an AI system sends to a customer on the firm's behalf is still a communication sent.

Insurance: the bulletin asks for traceability by name

The NAIC model bulletin expects a written AI systems program that addresses data lineage, traceability and auditability, and it lists data and record retention as an element of that program. It also says regulators may ask for documentation about data source, provenance and lineage for a specific AI system, and it explicitly includes generative AI. Our NAIC AI model bulletin checklist covers the full evidence file.

Healthcare: audit controls apply wherever ePHI flows

The HIPAA Security Rule's audit controls standard, 45 CFR 164.312(b), requires mechanisms that record and examine activity in information systems that contain or use ePHI. An LLM that reads a chart, drafts a note or answers a patient message is such a system. The rule does not prescribe a log retention period. The six-year rule in 164.316(b)(2)(i) applies to the documentation the Security Rule requires, and many organizations keep audit logs for the same period as a conservative choice. Our HIPAA-compliant AI development checklist covers the rest of the control set.

The EU AI Act, if you sell into Europe

Under Article 12, high-risk AI systems must technically allow automatic recording of events over their lifetime. Providers must keep the logs under their control for at least six months under Article 19, and deployers must do the same under Article 26(6), unless other law sets a different period. Following the 2026 Digital Omnibus amendments, the reference site lists these obligations as applying from December 2, 2027 for Annex III systems and August 2, 2028 for Annex I systems. Several widely shared articles still say August 2026; that date is out of date for high-risk obligations.

What to log for every LLM call: a field-by-field specification

The specification below is the core of this guide. It is organized into the five layers shown in the figure, and each field is tagged with how strongly we recommend it. "Required" means we would treat a system as non-compliant without it if it influences a regulated or customer-facing decision. "Recommended" means you will regret not having it. "Conditional" depends on the use case.

Diagram of an AI decision record organized into five layers linked by a trace ID: who and when, what went in, what the model was, what came out and what happened, and proof the record is intact
Five layers of one decision record. The trace ID ties each layer together and links every step of a multi-step agent run.

Layer 1: who and when

Field Level Why it matters
trace_id and parent_id Required Links every model call, retrieval and tool call in one run, and links runs to a case or ticket
timestamp_utc Required From a synchronized clock; examiners will line it up with other systems
human_user_id and role_at_time Required The person on whose behalf the AI acted, and their permissions at that moment
agent_or_service_identity Required Which agent, workflow or integration ran, distinct from the human
use_case_id Required Ties the event to your AI inventory entry, owner and risk tier
channel and tenant_id Recommended Web chat, email, API, internal tool; which customer or business unit
subject_id Conditional The applicant, patient, policyholder or customer the decision is about, as a pseudonymous key

The most common failure is a shared service account. If every call goes to the model under one API key with no human identity attached, you cannot answer "who did this," which is the first question in any investigation.

Layer 2: what went in

Field Level Why it matters
input_text or input_ref Required The exact prompt the user or upstream system sent, after redaction of secrets
system_prompt_version Required A version ID or hash; store the prompt text once in a versioned repository
retrieved_documents Required for RAG Document IDs, version or hash and chunk IDs, so you can show what the model read
records_read Required for agents Which customer or business records were loaded into context
policy_or_rule_version Recommended The underwriting guideline, credit policy or clinical protocol version in force
input_classification Recommended Data class detected (PHI, NPI, card data) and whether redaction ran

Retrieved documents are the field teams most often skip, and they are the field that explains most surprising answers. A retrieval-augmented system that answered correctly in March and wrongly in June usually changed because a source document changed, not because the model did.

Layer 3: what the model was

Field Level Why it matters
provider and model_requested Required What you asked for
model_returned Required The exact version string the provider reports back, which can differ from an alias
parameters Recommended Temperature, maximum tokens, tool choice and other settings
guardrail_results Required where used Content filters, PII detectors and policy checks, with pass, block or modify outcomes
prompt_template_version Recommended The application code or template release that built the prompt
response_id Recommended The provider's ID, for matching against their records if needed

Model aliases are a quiet source of audit gaps. If your code asks for a model family alias, the provider can move it to a newer version, and your behavior changes without a deployment on your side. Logging the returned version is how you prove which model produced a given output.

Layer 4: what came out and what happened

Field Level Why it matters
output_text or output_ref Required The exact response, including any structured output
finish_reason Recommended Normal stop, length cut-off, content filter or tool call
confidence_or_score Conditional Only if your system produces one; never as a substitute for reasons
decision_factors Required for consequential decisions The principal reasons, in terms you could put in a customer notice
human_review Required where a human is in the loop Reviewer ID, action (approved, edited, rejected), timestamp and any edits
final_outcome Required What the business actually did, which may differ from the AI's suggestion
tokens_and_cost Recommended Input and output tokens; useful for cost control as well as audit

Layer 5: proof the record is intact

Field Level Why it matters
record_hash and previous_hash Recommended Hash chaining makes deletion or editing of earlier records detectable
retention_class and retain_until Required Which retention schedule applies, set at write time
legal_hold Conditional Suspends deletion when litigation or an investigation is pending

Key takeaway: The two fields that most often decide whether a decision can be explained are the retrieved documents with their versions and the identity of the human who approved the outcome. Build those first.

What to log for AI agent decisions

An agent is different from a chat window: it plans, calls tools and changes records. Every one of those steps needs its own record, linked to the run by the trace ID. FINRA's list of agent risks maps neatly onto the fields you need. Acting without human validation calls for approval records. Acting beyond intended scope calls for authorization decisions. Multi-step reasoning that is hard to trace calls for step-level records.

For each step in an agent run, add these fields to the layers above:

  • Step number and type: plan, model call, retrieval, tool call, approval request, final answer.
  • Tool name and call ID: which function, API or MCP server was invoked.
  • Tool arguments: the parameters the agent passed, with secrets redacted.
  • Authorization decision: which permission or policy was checked, by which enforcement point, and whether it allowed or denied the call.
  • Tool result summary: status, records affected and a reference to the full response.
  • State change: the record IDs created, updated or deleted, with before and after values for high-impact fields.
  • Approval gate: whether a human approval was required, who gave it and what they saw when they approved.
  • Stop reason: completed, hit a step limit, blocked by policy, escalated to a human or failed.

Two design choices make this workable. First, log one structured record per step, not one blob per run, so you can query "every refund over $500 an agent issued last quarter" without parsing transcripts. Second, write the authorization decision even when the call is denied. Denied attempts are often the most interesting records in an investigation, because they show the agent trying to do something it should not.

If your agents connect to business systems through MCP servers, the server is a natural enforcement and logging point, since every tool call passes through it. We cover the build and maintenance side in MCP server development cost, and the governance agents need before production in from AI pilot to production.

Should you log the model's chain of thought?

Log it if your platform exposes it, but do not rely on it as the explanation. Anthropic's research on reasoning faithfulness, published in April 2025, planted hints in evaluation questions and checked whether models admitted using them. Claude 3.7 Sonnet mentioned the hint about 25% of the time on average and DeepSeek R1 about 39% of the time; most answers that relied on a hint did not disclose it. A model's stated reasoning is useful debugging evidence. It is not a reliable account of why the output came out the way it did.

Your explanation of record should therefore rest on things you can verify independently: the inputs, the retrieved documents and their versions, the business rules applied, the tool results and the human decision. Design the system so those factors are produced explicitly, for example by having the model return structured reasons that a rules layer checks, rather than reconstructing them from free-text reasoning afterwards.

Industry add-ons: banking, insurance, healthcare and broker-dealers

The five-layer record is the common core. Each regulated industry adds a few fields that examiners in that sector will ask for.

Industry Add these fields The question they answer
Lending and banking Principal reasons in adverse-action language; the credit policy version; any fair lending test run against the model version; link to the AI inventory entry and validation status Can you give the applicant specific reasons and show the model was governed?
Insurance Regulated practice affected (underwriting, rating, claims, fraud, marketing); data sources with lineage; drift monitoring result for the model version Is the AI system inside your AIS program, and can you show its data lineage?
Healthcare Patient identifier (pseudonymous); clinician who reviewed or signed; whether the output entered the record; BAA reference for the model provider Who accessed ePHI, through which system, and who took clinical responsibility?
Broker-dealers Whether the output was a communication with the public; supervisory reviewer; archive reference in your books and records system Was the communication captured, reviewed and retained?
HR and employment Role and requisition; criteria applied; human decision-maker; notices given Can you show a person made the decision and what the AI contributed?

The banking and insurance rows overlap with documentation you may already keep for traditional models. If you have that machinery, extend it to generative systems rather than building a parallel process. The AI acceptable use policy template we published earlier includes the policy clauses that make these logging duties enforceable for staff-facing tools.

How long to keep AI audit logs

Retention is where most AI logging plans fail quietly. Engineering teams default to 30 or 90 days in their observability tool because that is what the tool includes. Regulators measure in years.

Bar chart of minimum retention periods in months: EU AI Act high-risk logs 6, Regulation B consumer credit 25, SEC 17a-4(b)(4) communications 36, Colorado SB 26-189 compliance records 36, HIPAA documentation 72
Retention floors for specific record types, from the cited regulations as of October 2026. Your own schedule may need to be longer.

Three rules make retention manageable.

  1. Classify at write time. Every record gets a retention_class when it is written, derived from the use case. A customer service draft and a credit decision should not share a schedule.
  2. Take the longest applicable period. If an AI record supports a credit decision that is also a communication with the customer, the longer of the two schedules applies. Your records manager should own this mapping, not the engineering team.
  3. Separate hot and cold storage. Keep 30 to 90 days in your observability tool for debugging, and send the audit copy to cheaper write-once object storage for the long tail. Retrieval from archive can be slow; that is fine for an examination, as long as you have rehearsed it.

Storage volume is rarely the real cost. Text compresses well, and most of the expense of AI logging sits in the engineering time to capture the right fields, which is why it is cheaper to design in on day one than to retrofit. If you want to understand where AI run costs actually go, our guide to LLM API cost optimization breaks them down.

How to make an AI audit trail tamper-evident

An audit log that an administrator can quietly edit proves very little. NIST control AU-9, Protection of Audit Information, asks you to protect audit information and logging tools from unauthorized access, modification and deletion, and lists enhancements such as write-once media, storage on separate systems, cryptographic protection and dual authorization. For AI logs, four practices cover most of it.

  • Write-once storage. Cloud object stores offer write-once, read-many retention. On AWS, S3 Object Lock in compliance mode prevents any user, including the account's root user, from overwriting or deleting a protected object version or shortening its retention period. AWS states it has been assessed by Cohasset Associates for use in environments subject to SEC 17a-4, CFTC and FINRA rules. Azure and Google Cloud offer comparable immutability features.
  • Separate account and separate admins. Ship audit records to a storage account that the application team cannot administer. The people whose actions are logged should not control the log.
  • Hash chaining. Include the hash of the previous record in each new record, and periodically anchor a summary hash somewhere independent. A missing or edited record then breaks the chain.
  • Clock discipline. Use synchronized UTC clocks across the application, the model gateway and the tool services, or your sequence of events will not line up under scrutiny.

Key takeaway: Ask who could delete yesterday's AI logs without anyone noticing. If the answer is anyone on the application team, fix that before you add more fields.

What not to put in your AI audit logs

Logging everything is its own compliance problem. An audit store full of raw prompts becomes a second copy of your most sensitive data, with its own breach and access-request exposure.

Secrets never go in. API keys, bearer tokens, passwords and session cookies sometimes end up in prompts and tool arguments. Redact them before the record is written, not after.

Sensitive data goes in deliberately. Where a regulated decision requires the full input, store it in the restricted audit store with the same access controls as the source system. Use pseudonymous subject IDs in the general telemetry so engineers can debug without seeing names, diagnoses or account numbers. HIPAA's minimum necessary principle and state privacy laws both push in this direction.

Default observability will not capture content. This is the trap that catches many teams. The OpenTelemetry semantic conventions for generative AI, which most AI observability tools follow, mark the conventions as in development and state that instrumentations should not capture instructions, inputs and outputs by default. Tool call arguments and results are also opt-in. That is a sensible privacy default for telemetry, and it means a team that "turned on tracing" may have model names and token counts but no record of what was asked or answered. You have to enable content capture deliberately, route it to the right store and govern it.

Zero data retention is a provider setting, not your policy. Some providers offer zero data retention so they do not keep your prompts. That is often the right choice for privacy, and it means your own audit store is the only copy. Plan for that.

The AI audit trail reconstruction test

This is the most useful check we know, and it takes an afternoon. Pick one real decision an AI system influenced last month, ideally an adverse or unusual one, and hand the following ten tasks to someone who did not build the system. Give them read access to your logs and nothing else.

  1. Identify the human who initiated the request, and their role at that time.
  2. Identify the agent or workflow that ran, and its owner in your AI inventory.
  3. Produce the exact input, after redaction, and the system prompt version.
  4. List every document or record the system retrieved, with versions.
  5. Name the exact model version that produced the output.
  6. Produce every tool call, its arguments, the authorization decision and its result.
  7. Produce the output and any guardrail results.
  8. Show who reviewed or approved it, what they changed and when.
  9. State the principal reasons for the outcome in language you could send to the affected person.
  10. Show that none of these records could have been edited or deleted since.

Score each task pass or fail. Any failure on tasks 1, 3, 4, 8 or 9 is serious for a system that touches customers or regulated decisions. The failures this test most often exposes are shared service accounts (task 1), missing retrieval records (task 4), approvals that happen in email or chat with no link back to the trace (task 8) and logs kept only in a vendor dashboard with a short retention window (task 10).

Run it again whenever you change the model, the prompt template or the tools an agent can call. The test takes minutes once it passes, and it is the closest thing to a rehearsal for an examination.

Common AI audit trail mistakes and how to avoid them

Logging only the final answer. The final answer rarely explains itself. The retrieval, the tool results and the approvals do. Capture every step.

Relying on the model provider's logs. Provider logs exist for abuse monitoring, billing and debugging. Their retention is set by the provider, they cannot see your users, documents or tools, and they may hold nothing under zero data retention. Keep your own trail.

Approvals outside the system. A manager approving an agent's refund in a chat thread leaves no link between the approval and the action. Build approval into the workflow so it lands on the trace.

Unpinned models. If you call a model alias and log nothing about the version returned, you cannot prove which model made a decision after the provider updates the alias.

No owner. Logging specifications decay when nobody owns them. Name an owner for the standard and an owner for each AI system in your inventory, and review the standard when a rule changes. Several changed this year.

Retrofitting at the end. Adding identity propagation, retrieval capture and approval records after launch means touching every component. It is far cheaper to make the decision record part of the design. Our guide to how long it takes to build an AI agent shows where this work sits in a realistic schedule.

Conclusion: build the decision record first

AI audit trail requirements in the US are not one rule but a convergence. HIPAA, Regulation B, broker-dealer recordkeeping and the NAIC bulletin already reach the records AI systems create. Colorado's SB 26-189 and California's ADMT regulations add explicit explanation and record duties starting in 2027, and the EU AI Act sets logging rules for high-risk systems sold into Europe. Each of them is easier to satisfy with one well-designed decision record than with a scramble when the first request arrives.

The practical sequence is short. Inventory the AI systems that touch customers or regulated decisions. Adopt the five-layer record, starting with identity, retrieved documents and approvals. Assign a retention class at write time and send the audit copy to write-once storage your application team cannot administer. Then run the reconstruction test on one real decision and fix what fails.

Fleurant AI's AI compliance team builds this kind of logging, evaluation and control layer for companies in banking, insurance and healthcare, alongside the AI agents themselves. If you would like a second opinion on an existing system or a logging standard you are drafting, talk to a specialist, or start from the Fleurant AI home page to see how the pieces fit. For questions that turn on legal interpretation, your counsel and compliance team should make the final call.

Frequently asked questions

Is there a single law that sets AI audit trail requirements in the US?

No. As of October 2026 there is no federal statute that defines an AI audit log. The requirements come from rules that already govern your industry, such as the HIPAA Security Rule's audit controls, Regulation B's adverse action and record retention rules, SEC Rule 17a-4 for broker-dealer communications and the NAIC model bulletin for insurers, plus newer state laws such as Colorado SB 26-189 and California's ADMT regulations. The audit trail is how you prove compliance with all of them.

Should we log full prompts and responses?

For any AI system that influences a regulated or customer-facing decision, usually yes, because without the exact input and output you cannot reconstruct or explain the decision. Store them in a restricted, encrypted, write-once store separate from operational telemetry, redact secrets before writing, and apply the same access controls you use for the underlying records. For low-risk internal tools, metadata plus a sampled content log is often enough.

How long should we keep AI audit logs?

Match the longest rule that applies to the record the AI produced or influenced. Floors we found include six months for high-risk system logs under the EU AI Act, 25 months for consumer credit application records under Regulation B, three years for broker-dealer communications under SEC Rule 17a-4(b)(4), three years for compliance records under Colorado SB 26-189 and six years for HIPAA-required documentation. Confirm the final schedule with your counsel and records manager.

Do we need to log an AI model's chain of thought?

Log the reasoning or plan text when your platform exposes it, because it helps engineers debug and helps reviewers see what the agent intended. Do not treat it as the explanation of record. Anthropic's 2025 research found reasoning models often fail to mention information they actually relied on. Your explanation should rest on the inputs, retrieved documents, rules applied and tool results, which you can verify.

Are our model provider's logs enough?

Rarely. Provider logs are built for the provider's purposes: abuse monitoring, billing and debugging. Retention is set by their policy, not your regulator, and zero data retention options mean the provider may hold nothing at all. They also cannot see your user identities, retrieved documents, tool calls or human approvals. Keep your own trail and treat the provider's logs as a supplement.

What is the difference between AI observability and an AI audit trail?

Observability answers engineering questions in near real time: latency, cost, errors and quality. An audit trail answers accountability questions months or years later: who asked, what the system saw, what it decided, who approved it and whether the record is intact. They can share a pipeline, but audit records need identity, retention, write-once storage and access controls that observability tools do not provide by default.

Sources

  1. SB26-189: Automated Decision-Making Technology, Colorado General Assembly
  2. California Finalizes Regulations to Strengthen Consumers' Privacy (September 23, 2025), California Privacy Protection Agency
  3. Article 12: Record-Keeping, EU Artificial Intelligence Act (artificialintelligenceact.eu)
  4. Article 19: Automatically Generated Logs, EU Artificial Intelligence Act (artificialintelligenceact.eu)
  5. Article 26: Obligations of Deployers of High-Risk AI Systems, EU Artificial Intelligence Act (artificialintelligenceact.eu)
  6. 45 CFR 164.312: Technical safeguards, Legal Information Institute, Cornell Law School
  7. 45 CFR 164.316: Policies and procedures and documentation requirements, Legal Information Institute, Cornell Law School
  8. 12 CFR 1002.9: Notifications (Regulation B), Legal Information Institute, Cornell Law School
  9. 12 CFR 1002.12: Record retention (Regulation B), Legal Information Institute, Cornell Law School
  10. 17 CFR 240.17a-4: Records to be preserved by certain exchange members, brokers and dealers, Legal Information Institute, Cornell Law School
  11. 2026 FINRA Annual Regulatory Oversight Report: GenAI, FINRA
  12. NAIC Model Bulletin: Use of Artificial Intelligence Systems by Insurers (adopted December 4, 2023), National Association of Insurance Commissioners
  13. SR 26-2: Revised Guidance on Model Risk Management, Board of Governors of the Federal Reserve System
  14. NIST SP 800-53 Rev. 5, AU-3: Content of Audit Records, CSF Tools (NIST SP 800-53 reference)
  15. NIST SP 800-53 Rev. 5, AU-9: Protection of Audit Information, CSF Tools (NIST SP 800-53 reference)
  16. Semantic conventions for generative AI spans, OpenTelemetry
  17. Reasoning models don't always say what they think, Anthropic
  18. Locking objects with Object Lock, Amazon Web Services

Free, no-obligation consultation

Have a question about AI audit trails?

Tell us what you're working on or what you'd like to know. A specialist will get back to you with practical next steps, whether or not we end up working together.

  1. 1Send your question or project details (takes 2 minutes)
  2. 2A specialist reviews it and replies within 1 business day
  3. 3Get clear, practical next steps, free