Consider a composite situation, not a real engagement. A controller opens the monthly bank reconciliation, and the matching tool shows 1,250 unmatched items out of 18,400 bank lines. The accounting manager reviewed every item and approved the package. The external auditor asks which version of the matching logic produced that list, whether the bank file was complete when the tool ran, and what the reviewer saw on screen. Nobody can answer. The matching feature arrived in an ERP update that the supplier switched on, and the SOX risk and control matrix still describes a spreadsheet.
If AI influences financial reporting, your organization must be able to show five things: where the AI is used, how it works, that its inputs and outputs are reliable, that changes to it are controlled, and that a human exercised oversight and left evidence of it. Each of those five becomes an auditor question, and this article takes them in the order a SOX program would address them. Treat the five questions as a practical synthesis of current auditor concerns. No official PCAOB or SEC checklist with five questions exists, and I would be wary of anyone who presents one. The article is written for controllers, SOX control owners, and internal auditors who need working answers before fieldwork starts.
You do not need a separate AI compliance program to get there. COSO released guidance on internal control over generative AI in February 2026 that applies its Internal Control Integrated Framework to generative AI, and it names the risks that matter most to financial reporting: prompt-based manipulation, opaque reasoning, model drift, and frequent configuration changes. SEC staff remarks at the December 2025 AICPA and CIMA conference, as summarized in an EY compendium, pointed the same way, with attention on model design, data, human oversight, and the way AI changes IT general controls. What follows turns that direction into an inventory, a control narrative standard, a validation plan, a change register, and a failure path you can build this quarter.
When Does AI Become a SOX Matter?
AI becomes a SOX matter when it influences information, decisions, evidence, or actions that support internal control over financial reporting. The name of the tool does not decide scope. Neither does the vendor's marketing. A model that drafts a meeting agenda sits outside the financial reporting process. A model that proposes account coding for invoices, flags journal entries, matches bank lines, or drafts a disclosure paragraph sits inside it, and an auditor will ask how much anyone relies on its output. The useful question is therefore never whether the company uses AI. It is where AI influences a financial reporting process, a control decision, an accounting estimate, an evidence population, or a disclosure.
Classify each use by the role the AI plays, because the control response differs for each role. Productivity support with no effect on financial reporting usually needs a usage policy and nothing more. An input or processing component that supports an existing control belongs in that control's description and evidence. A control activity that directly detects or prevents a financial reporting error needs a design of its own, plus testing and monitoring on a schedule. A system component whose failure could cause a material misstatement needs the full set of IT general controls around it. An agentic workflow that can take actions or change records needs all of that plus hard limits on what it may do without a human.
The classification rests on potential financial reporting impact. Take an AI recommendation that the controller approves. The approval looks like a normal human control, yet the recommendation shaped what the controller saw, and the matrix that omits the AI step describes a process that no longer exists. That omission is the scoping gap I expect auditors to find most often. Write the AI step into the control description, name its owner, and state what the human relies on.
Three different uses of AI also need separating, because the control implications differ. The first is AI inside the company's financial reporting process. The second is AI that internal audit or the SOX team uses to test controls. The third is AI that the external auditor uses on the audit. The PCAOB staff Spotlight on generative AI describes outreach in which audit firms used it mainly for administrative and research work, while preparers used it for drafting internal documents and for less complex tasks such as preparing account reconciliations or identifying reconciling items. That Spotlight carries the views of PCAOB staff, not the Board, and it dates from mid-2024, so read its picture of adoption as a floor. It also describes a PCAOB research project on whether guidance or standard changes are needed, and you should check where that project stands before you tell an audit committee what the standards require.
Which Five Characteristics of GenAI Change How You Design Controls?
COSO builds its GenAI guidance on its earlier work for robotic process automation, and the difference between the two explains most of what follows. RPA bots follow static rules and do the same thing every time, so control design could focus on whether the bot executed correctly. Generative AI reasons probabilistically and can adapt as it goes, and it creates new content, which extends its influence beyond carrying out tasks into shaping analyses and judgments that feed decisions. COSO states five foundational characteristics that guide how controls should be built or adapted, whatever the use case or technology. Each one has a direct SOX consequence.
The first characteristic is that GenAI is probabilistic, not deterministic, so its outputs can be confidently wrong. COSO's answer is to treat outputs as claims that need validation, and to stop accepting them as facts by default. For a SOX program this means an AI output on a screen is not evidence of anything until someone has checked it against a source. A control description must say what supports the claim, who checks it, and what happens when the check fails. If management relies on AI output as evidence that a control operated, an auditor will ask what independent support exists.
The second characteristic is that GenAI is dynamic. Models and prompts can change often, and so can the data used for retrieval, which is why COSO calls for risk assessment and change control that run continuously or move in that direction, with monitoring to match. An annual walkthrough cannot keep pace with a system whose behavior shifts between two month-end closes. The practical response is a version record for every component that shapes the output, plus a trigger list that forces reassessment when any of them changes. Question four below builds that record.
The third characteristic is that GenAI scales easily, for better and for worse. Automation can multiply quality and efficiency, and it can multiply errors and bias with equal speed. Suppose, as an assumption, that a classification model has a one percent error rate on 18,400 lines each month. That is 184 wrong lines every close, and they all share a cause, so a single fix repairs the whole batch. Controls should stop small errors from compounding into systemic issues, which in practice means batch-level thresholds that hold a population before it reaches the ledger.
The fourth characteristic is the low barrier to entry. Almost any analyst with a browser can build or use an AI tool, so COSO directs controls at who may build or deploy GenAI and who may interact with it. In finance this is the shadow AI problem: an extraction script, a reconciliation assistant, or a reporting summary that nobody registered. Access rules and an inventory that finance owns are the response, and the next section describes how to build that inventory.
The fifth characteristic cuts the other way: GenAI can help govern GenAI. COSO points to automated multi-model validation, where independent models cross-check outputs to catch inconsistency and drift, including bias, and it notes that this can make monitoring and validation practical at a scale people cannot match. I think this is right, with one condition. The validating model is itself an AI component, so it belongs in the inventory, it needs its own change control, and it must be independent enough that both models do not fail in the same way. Expect the auditor to ask who validates the validator.
Question 1: Where Does AI Influence Financial Reporting?
The auditor wants a process-level answer, and a statement that the company uses AI does not qualify. Map each AI use to the financial statement accounts it touches, the assertions at risk (completeness, existence, accuracy, valuation, rights, and presentation), and the process it sits in: journal entries, account reconciliations, revenue, close and consolidation, significant estimates, management review controls, disclosures, or the production of audit evidence. The failure mode here is a scoping gap. Most organizations keep an inventory of applications, and almost none keep an inventory of the AI-enabled features inside finance, ERP, workflow, analytics, forecasting, and reporting tools.
The gap opens in five predictable ways. A finance assistant extracts or classifies accounting data and nobody registered it. A supplier adds an AI feature to an application that was already in scope, and the release note went to an IT mailbox. Employees use public or enterprise AI tools to prepare reporting support. A human approves an AI recommendation, so the process looks manual even though the AI shaped the result. The SOX matrix then describes the human review and leaves out the step that generated what the human reviewed.
COSO offers a useful interview aid in the form of eight capability types, arranged in the order data moves toward a decision. They are data extraction and ingestion, data transformation and integration, automated transaction processing and reconciliation, workflow orchestration and autonomous task execution, judgment, forecasting and insight generation, AI-powered monitoring and continuous review, knowledge retrieval and summarization, and human and AI collaboration. The lens looks at what the AI does and ignores the product name, which keeps it stable when vendors rebrand. Ask each process owner which of the eight appears in their close steps. COSO's own example of workflow orchestration is an agent that pulls trial balance and subledger data, runs reconciliations, assigns follow-up tasks, escalates exceptions, and compiles a reconciliation package, and that single example touches four capability types at once.
Record what you find in an ICFR-relevant AI inventory. The table shows the fields I would require, with an invoice classification feature as the example.
| Field | What to record | Example |
|---|---|---|
| AI system or feature | Name and function | Invoice classification model |
| Business process and owner | Process, accountable role | Accounts payable, corporate controller |
| Supplier or builder | Third party or internal | ERP provider |
| Financial statement impact | Accounts and assertions | Expense classification, accrual completeness |
| Human decision point | Who decides and when | Controller reviews exceptions above threshold |
| Data sources | Systems and populations | Invoice records, purchase orders, vendor master |
| Version and effective date | Model, prompt, configuration | Version identifier with date |
| Evidence retained | What an auditor can request | Input population, outputs, review, exceptions |
The inventory must include embedded AI, third-party functionality, internally built models, and known shadow use, and it must be reconciled to the risk and control matrix each quarter and after any material change. COSO's guidance emphasizes ongoing inventories of AI use cases, defined ownership, and escalation paths, which matches this approach. Reconciliation is where the value sits: every SOX-relevant AI use should map to a matrix entry, and every matrix entry that relies on AI should map back to an inventory line.
The warning signs are easy to spot. IT built the inventory and finance never saw it. Supplier release notes have no reader in the controller's team. A reviewer cannot say which step of the process the AI performs. My decision rule is simple: if nobody can name the owner and the assertion for an AI use, treat it as in scope until someone proves otherwise, and block it from finance data in the meantime. Evidence to keep includes the dated inventory export, the reconciliation to the matrix, and an attestation from each process owner that the inventory is complete for their process.
Question 2: Can You Explain How the AI-Enabled Control Works?
This question tests control design and understandability, and it is answerable by people who are not data scientists. The auditor needs to understand what the AI receives, what processing it performs, what it produces, which financial reporting risk it addresses, who reviews the output, what counts as an acceptable result, what happens when the output is incomplete or wrong, and what evidence remains. Vague descriptions fail. Specific ones pass. A control that says management reviews AI output names no input, no threshold, no reviewer, and no evidence, and a statement that the AI platform validates financial data is no better.
A stronger description reads like an operating procedure. Each month, the system compares the complete population of posted journal entries against documented risk indicators. The controller reviews every entry that exceeds the documented threshold, investigates each exception, records a disposition, and keeps the input population, the model version, the output, the review evidence, and the remediation record. Every noun in that description is something an auditor can request and inspect. Operational language beats technology language here, and I would not sign off on a description that leans on words like intelligent or automated validation.
Write the control as a chain: source data, data preparation, model or rule execution, output, human review or automated action, exception handling, and evidence retention. For each link, document the purpose, the owner, the system, the inputs and transformations, the decision criteria, the output, the review procedure, the escalation path, and the evidence kept. Produce a plain-language narrative for the chain and support it with a technical appendix for the people who need depth. Then walk one real or representative transaction through every link, because an auditor will do the same during a walkthrough.
A reconciliation example shows the level of detail required. The matching model compares the general ledger balance with subledger records each month and identifies mismatches using configured matching logic. The accounting manager reviews every unmatched item, confirms the cause against source documentation, and records whether the item was corrected, accepted with justification, or escalated. The control does not post adjustments automatically. That last sentence matters most, because it fixes the boundary between AI analysis and human judgment, and an auditor reading it knows exactly where responsibility sits.
Define that boundary for every use. State whether the AI may recommend, hold, route, or post, and state who approves each action. If the AI can post entries and the person who approves the output is the same person who configured the AI, the segregation of duties problem is obvious, and an agentic workflow makes it larger. Watch for narratives that a supplier wrote and nobody on the finance team can restate. Watch also for narratives that describe the intended process while the team works differently, a gap that I return to in the discussion of question five.
Question 3: How Do You Know the AI Is Reliable and Operating as Designed?
You know it by connecting reliability to the financial reporting risk, validating before deployment, and monitoring after it. A model that gives consistent or plausible answers is not thereby reliable, because a plausible answer built from incomplete data is the most dangerous kind. The auditor will ask what validation happened before go-live, which known outcomes were used, what error thresholds apply, how false positives and false negatives are evaluated, what the model does when it is uncertain, whether it was tested on the complete population, and how performance is tracked over time. Each answer needs evidence, and the evidence needs an owner.
Start with the inputs, because AI output cannot be more reliable than the data supplied to it. A model can produce a convincing result from data that is incomplete, stale, duplicated, mapped to the wrong account, or pulled without authorization. This is the reliability of information produced by the entity, and an auditor will ask whether the source population is complete and accurate, whether the data transformations were validated, whether excluded records are identified, whether interfaces and APIs operate correctly, and whether you can reproduce the exact input behind a prior decision. Prompt instructions and retrieval sources count as inputs too, since a changed instruction can change the result. Completeness and accuracy checks are part of the control evidence, and the team that treats them as technical prerequisites will have nothing to show.
The PCAOB staff Spotlight supports this emphasis. Audit firms described the importance of auditability for both the source data and the AI-created content, and some designed their tools to record the underlying source data used. Preparers pointed to the black-box nature of some tools and to inconsistent output as reasons the auditability of AI-generated content is in question, and some were running tools in a parallel environment before deployment. A parallel run is a good model for validation: operate the AI beside the existing control for several closes and compare the results. Keep the comparison.
Monitoring after deployment needs metrics tied to the control objective. Useful ones include error rate, false-positive rate, false-negative rate, exception volume, override rate, unprocessed population, and data reconciliation failures. A metric works only if somebody knows what to do when it breaches its threshold, so every metric needs an owner and a documented action. Consider an illustrative scenario with assumed numbers, not an actual engagement. A matching model processes 18,400 bank lines, auto-matches 17,150 of them (93.2 percent), and routes 1,250 to a person. Each month the controller's team reperforms a random sample of 400 auto-matched lines, and the error rate in that sample ran at 0.5, 0.7, and 0.6 percent over the previous three months. This month six of the 400 are wrong, which is 1.5 percent and above the assumed tolerance of one percent, so the documented action is to hold the next batch and investigate before anything posts. The controller owns the metric. The sample workpaper is the data source, and the evidence is the sample selection and the reperformance, plus the recorded hold decision.
Drift and vendor change need the same discipline. SEC staff remarks reported from the December 2025 conference said auditors should adjust their IT general controls approach for AI models, because model drift and third-party changes call for more proactive and frequent monitoring by both the registrant and the auditor. Annual testing alone is inadequate for a system that changes often or processes large populations, so add continuous monitoring where the risk justifies it. Preserve the input, the model or configuration, the output, and the timestamp for every run, because reproducibility is what lets an auditor evaluate the control independently. Finally, define a failure procedure for when the AI is unavailable or unreliable, since question five depends on it.
Question 4: Who Can Change the Model, Prompts, Data, or Rules?
This question tests IT general controls, access, segregation of duties, and change management, and the answer must reach well beyond source code. AI behavior changes through many doors that a conventional release process never watches. Identify who can change each of the following, and treat every one as controlled configuration. The model and its version come first, followed by the system prompt and the prompt templates users run. Retrieval documents and any training or fine-tuning data also shape the output, as do decision thresholds and the business rules wrapped around model output. Add the API connections and the tools an agent can call, the user permissions (including those granted to the AI itself), and the human-review requirements. Finally, list every configuration that the supplier manages on your behalf.
A defensible change record shows what changed, who requested it, who approved it, why it changed, what risk assessment was done, what testing was completed, which version became effective, and whether the change required revalidation. It also shows whether the SOX control description or the evidence requirements changed as a result. Many teams record the first five items and forget the last two, which means the matrix drifts away from the system. I would add one rule: a change to a prompt in a SOX-relevant use follows the same path as a change to a posting rule, because both alter what reaches the ledger.
Access should be role-based and reviewed periodically, and duties should be separated where practical. The person who develops or changes the AI should not automatically be the person who approves its production use or evaluates its effectiveness for financial reporting. The PCAOB staff Spotlight records the view of some audit firms that AI use by preparers could amplify existing IT risks, segregation of duties among them, or create new ones. SEC staff remarks also pointed to evolving fraud risk, including data poisoning and prompt engineering. Both concerns land on the same control: whoever can edit the instructions or the data can influence the output.
Third-party AI needs its own procedure, because the supplier can change the model without your release process ever running. Obtain supplier documentation, contractual commitments on change notification, testing information, and a SOC 1 report or other assurance where applicable. Read the complementary user entity controls in any SOC 1 report and confirm that someone at your company performs them. A supplier report does not remove your responsibility to assess how the service affects your own internal control over financial reporting. Set a trigger so that a major supplier update starts a reassessment within a defined number of days, and record who owns that clock.
Question 5: What Happens When the AI Is Wrong, Uncertain, or Changed?
You need a documented failure path that exists before the first incident. The auditor may ask for a specific example: the last material or unusual exception, how it was investigated, who investigated it, whether financial reporting was affected, how the AI was corrected or bypassed, whether management or the audit committee heard about it, and what prevents a repeat. A team that can only describe what would happen in theory has a design gap. A team that can produce the last real exception with its full record has a control that operates.
Design the path as a sequence with named owners at each step:
- The AI shows uncertainty or abnormal output.
- The transaction or batch is stopped, held, or routed to a person.
- A human investigates.
- The team assesses the impact on financial reporting.
- The control is corrected or independently performed by a person.
- Root cause and evidence are recorded.
- Remediation is approved.
- The AI is retested before it returns to normal operation.
The fallback must be designed in advance. State who can suspend the AI, who can approve manual processing, how the backlog is handled when the tool is down at quarter end, and what evidence is preserved during the manual period. Test the switch once a year in a tabletop exercise, because a suspension procedure that nobody has rehearsed will fail on the day it matters.
An AI-detected anomaly is not automatically a control failure. A practitioner analysis on AI-enabled SOX and internal audit testing describes what it calls the inference gap: long-standing processes often rely on undocumented knowledge, so automated tools that evaluate only the data and configured logic can generate results that conflict with established practice. In its example, an automated test flagged transactions as missing approvals even though the approvals were in the evidence, and the real problem was an imprecisely written test procedure. The analysis addresses AI used by internal audit and SOX teams for testing, and the same logic applies to AI inside the reporting process. Its investigation questions are practical: what human assumptions are missing from the data or the logic, whether the source data is complete and accurate, whether the anomaly is a true control failure or evidence that information produced by the entity lacks the precision automated governance needs, and what evidence should be kept from the investigation.
Ownership matters in that investigation. The analysis argues that data remediation belongs with the business owner and not with the internal audit team, since the first line owns its data. I agree, and I would add a reporting rule: record every confirmed AI-related exception in a log that goes to the controller and, when thresholds are crossed, to the audit committee. The log feeds your annual risk assessment and shows an auditor that exceptions lead somewhere. Without it, the same anomaly returns every quarter under a different ticket number.
How Do You Test Whether Human Review Is Real?
A control does not become human-in-the-loop because a person clicks approve. Shadow reliance is my term for the condition where the documented process says a human decides, and in practice the human routinely accepts the AI output without meaningful challenge. SEC staff remarks stressed maintaining human oversight and involvement in AI-driven internal controls. The PCAOB staff Spotlight reports that audit firms expect the person who uses an AI tool to remain responsible for the results and that supervisors apply the same diligence as for work done without AI. Preparers voiced the same expectation about reviewing AI output. An auditor testing your control will therefore test the reviewer as well as the review record.
Expect questions on eight attributes. Does the reviewer have sufficient competence for the subject? Does the reviewer receive the underlying evidence and not only the AI summary? Does the reviewer understand the limits of the AI? Does the reviewer challenge unusual results? Does the reviewer have authority to reject or override the output? Is the review documented? Does it happen before the output affects reporting? Does the reviewer escalate uncertainty and control exceptions? A yes on all eight needs support, and each yes has a different kind of evidence behind it.
You can measure several of these directly. Track the override rate and the time spent per item, along with the number of items each reviewer handles in a day. In my view, an override rate of zero over a full year is a warning sign, not a success metric, because it usually means nobody is looking. Internal audit can also seed a review population with known errors and see whether the reviewer finds them, a practice that tests attention without relying on self-reporting. Keep the results as evidence, and keep the reviewer's reasoning when they accept a flagged item.
Fix the design when the measures disappoint. Volume is the most common cause: 1,250 unmatched items at an assumed six minutes each is 125 hours of review in one close, and no reviewer sustains real attention over that load. Raise thresholds so that only material or unusual items reach a person, and route routine items to a sampling control. Give the reviewer a view that shows source documents beside the AI output. Train reviewers on how the model fails, since a person who knows the typical failure patterns catches them. Document the competence requirement in the control description so that reassigning the review to a new hire triggers training.
How to Close the Gaps With Seven Practical Steps
Start with the process, not the technology inventory. Walk through close, reporting, consolidation, revenue, purchasing, payroll, tax, treasury, and disclosure, and ask at each stage whether AI generates, transforms, classifies, prioritizes, reconciles, approves, or validates information, whether a human relies on the result, and whether the AI sits inside a supplier application. This is step one, and its output is a list of candidate uses. Step two classifies each candidate using the five roles from the first section, based on potential financial reporting impact and not on whether the tool carries an AI label. Uses with no effect on reporting leave the SOX program, and the rest continue.
Step three links each remaining use to the financial statement account, the relevant assertion, the plausible misstatement, the AI contribution, the human and automated controls, the control owner, and the evidence required. An invoice classifier could misclassify expenses or miss invoices, which affects accuracy and classification, along with completeness. The control response includes input-population reconciliation, model performance testing, exception review, and documented approval of corrections. Write the example down in the matrix, because an entry that names the plausible misstatement is far easier to test than one that names a tool. Owners should sign the entry, and the signature should carry a date.
Step four tests the complete control chain and avoids testing only the final approval screen. Test the source population, extraction, transformations, model or rule version, prompt or configuration, output, human review, exception resolution, the final accounting result, and evidence retention. Many AI controls fail at this step, because the organization can show an approval but cannot reconstruct what the reviewer saw or which version made the recommendation. Reperform one item from end to end before you rely on the control. If you cannot, the design has a hole, and no amount of sampling downstream will fill it.
Step five sets the evidence standard, which the next section lays out. Step six adds continuous monitoring where risk requires it, with metrics and thresholds tied to the control objective and an owner for each breach action. Step seven defines the triggers for reassessment: a new model, a major supplier update, a new data source, a new business process, a changed prompt or rule, a changed materiality level or threshold, a new autonomous capability, a significant performance decline, a control failure or material exception, and a change in how much the organization relies on the AI output. Each trigger should name who starts the review and within what period. Share the method with your external auditor early, since SEC staff remarks encouraged early and frequent communication between issuers and auditors as AI risk assessments become more dynamic.
Which Evidence Should You Retain for an AI-Enabled Control?
Define the minimum evidence for each control before the auditor asks, and retain it so that a third party can reproduce or independently evaluate the control. Good evidence is complete for the population, time-stamped, attributable to a named person or system, and stored where it cannot be edited after the fact. A screenshot of an approval screen meets none of those tests on its own. It shows that someone clicked, and it says nothing about what the screen contained.
Supplier-held evidence deserves a separate check. Many AI tools keep run logs for a short period, and that period may be shorter than your SOX retention requirement. Ask the supplier how long logs and version histories persist, whether you can export them, and whether an export carries enough detail to rebuild a past run. If the answer is no, schedule your own capture at each close, and document the capture as part of the control. The standard below applies across uses, and you can adapt the rows to each control. The owner column assumes a finance-led program with IT and the supplier supporting.
| Control element | Evidence to retain | Typical owner |
|---|---|---|
| Input population | Source totals, completeness check, extraction logic, data lineage | Process owner with IT |
| Model or application | Version, effective date, supplier release notes | AI system owner |
| Prompt or configuration | Version, change record, approver | AI system owner |
| Processing | Date and time, run log, population processed | IT operations |
| Output | Result, confidence or exception indicators | Process owner |
| Human review | Reviewer identity, date, procedure, challenge or override | Control owner |
| Exception handling | Disposition, root cause, escalation record | Controller |
| Validation | Test results, reperformance, drift metrics | SOX or internal audit |
| Change and access | Change register, access review results | IT governance |
One more point on retention: keep the evidence for AI-enabled controls in the same repository as your other SOX evidence, with the same access rules and the same owner for completeness. A parallel store that only the AI team can open invites exactly the reconstruction failure this section describes.
The point of the table is reconstruction. If the auditor selects one December item, you should be able to show the input population as it stood on that day, the version that processed it, the output, the reviewer's screen and decision, and the disposition of any exception. Retention periods follow your existing records policy for SOX evidence, and you should store the evidence in a location the audit team can reach without asking engineering for an export. Auditors tend to test the first row and the last row, because inputs and change records are where reconstruction breaks. The sample you pull for your own testing should follow the same logic, and a spreadsheet version of this table works well as the request list.
AI SOX Control Objectives and Audit Procedures
| # | Common topic / SOX control area | Control type | Control objective | Control audit procedure | Control evidence to assess effective control attributes |
|---|---|---|---|---|---|
| 1 | GenAI Governance & Accountability | Entity-level governance / management control | Establish accountable ownership, authority, and oversight for GenAI used across the organization, including systems, prompts, datasets, retrieval sources, configurations, and AI-enabled ICFR processes. | Inspect the GenAI governance framework, inventory, ownership assignments, RACI, escalation paths, and governance minutes. Confirm material AI-enabled ICFR processes have named owners with appropriate authority and accountability. | Approved AI governance policy; GenAI inventory; RACI; system owner assignments; governance committee minutes; escalation records; documented accountability for AI-enabled controls. |
| 2 | GenAI Acceptable Use & Ethical Boundaries | Entity-level preventive control | Ensure GenAI usage complies with organizational, legal, regulatory, privacy, confidentiality, intellectual-property, and ethical requirements, particularly for financial reporting activities. | Inspect the approved GenAI acceptable-use policy. Determine whether prohibited data, restricted use cases, required approvals, transparency expectations, and escalation requirements are defined. Sample user activity or attestations for compliance. | Approved AUP; prohibited-use matrix; data restrictions; legal/compliance approval; user acknowledgements; exception approvals; policy violation records. |
| 3 | GenAI Risk-Based Scoping | Risk assessment / SOX scoping control | Identify GenAI applications that could materially affect financial reporting, including systems producing, transforming, validating, reconciling, summarizing, or supporting financial information. | Obtain the GenAI inventory and compare it with the SOX application inventory, key controls, financial processes, and system dependencies. Challenge whether AI use cases affecting ICFR have been appropriately scoped. | AI inventory; SOX application inventory; process narratives; risk assessments; system dependency maps; documented scoping decisions; management sign-offs. |
| 4 | Use-Case Objectives & Suitability | Preventive management control | Ensure each GenAI use case has defined objectives, boundaries, success criteria, and appropriate technology selection before being deployed in an ICFR process. | Select high-risk GenAI use cases and inspect documented objectives, intended users, acceptable outputs, limitations, regulatory requirements, and rationale for using GenAI rather than deterministic automation or other technology. | Use-case approval; business requirements; risk/benefit assessment; success criteria; intended-use documentation; technology selection rationale. |
| 5 | GenAI Risk Assessment | Risk assessment control | Identify and assess GenAI risks that could prevent financial reporting objectives from being achieved, including hallucinations, bias, drift, data leakage, prompt injection, and third-party dependencies. | Inspect risk assessments for material AI-enabled processes. Verify risks address both inherent and residual risk, likelihood and impact, changing data, model updates, vendors, regulations, and downstream financial reporting consequences. | AI risk register; risk assessments; risk scoring; scenario analysis; identified mitigating controls; management approval; risk acceptance documentation. |
| 6 | AI Change Risk Assessment | Change-management risk control | Ensure significant changes to models, prompts, training data, retrieval sources, safety controls, integrations, or vendors trigger timely reassessment of financial reporting risks. | Sample significant AI changes during the audit period. Determine whether changes were identified, risk assessed, approved, tested, and evaluated for potential effects on ICFR before implementation. | Change tickets; impact assessments; risk reassessments; testing results; approval records; vendor notifications; post-implementation reviews. |
| 7 | AI Fraud Risk | Fraud risk assessment / preventive control | Identify and mitigate AI-enabled fraud risks, including synthetic records, deepfakes, manipulated outputs, crafted prompts, excessive agent authority, and unauthorized automated actions. | Inspect fraud risk assessments for AI-enabled financial processes. Evaluate whether management considered AI-specific fraud scenarios and whether preventive and detective controls address identified risks. | Fraud risk assessment; fraud scenarios; control mapping; investigation records; monitoring reports; management certification; remediation documentation. |
| 8 | AI Governance Competence | Entity-level / training control | Ensure personnel operating, reviewing, approving, or developing GenAI controls possess competence appropriate to their role, risk exposure, and responsibility. | Inspect training requirements and completion records. Sample control operators, reviewers, developers, and managers to determine whether training covers acceptable use, AI limitations, secure prompting, review responsibilities, and escalation. | Training curriculum; completion reports; role-based training matrix; certifications; competency assessments; retraining records. |
| 9 | AI Performance Accountability | Management performance control | Establish accountability for AI performance, safety, compliance, control adherence, and business outcomes, including consequences for repeated control violations or negligent oversight. | Inspect performance objectives and management reviews for selected AI-enabled processes. Determine whether AI performance and control adherence are monitored and whether repeated deficiencies result in corrective action or retraining. | Performance objectives; KPI/KRI dashboards; performance reviews; corrective-action records; retraining evidence; escalation documentation. |
| 10 | AI Access Management | ITGC logical access control | Restrict access to GenAI models, prompts, configurations, retrieval sources, datasets, integrations, and administrative functions based on business need and authorized roles. | Test user access for selected GenAI systems. Compare access with approved roles and responsibilities, investigate privileged access, and test timely removal or modification of inappropriate access. | User access listings; role matrices; access requests; approvals; privileged-access reports; termination reports; periodic access reviews. |
| 11 | Segregation of Duties | ITGC / preventive control | Prevent individuals from independently configuring AI systems and approving or relying on their own AI-generated outputs where segregation is required. | Identify users with configuration, deployment, approval, and review privileges. Assess incompatible access combinations and inspect compensating controls where segregation cannot be achieved. | SoD matrix; role definitions; access reports; conflict reports; compensating-control evidence; management review and approval. |
| 12 | AI Configuration Management | ITGC change management | Ensure prompts, system prompts, thresholds, retrieval connectors, transformation rules, model configurations, and other AI components are formally governed as configuration items. | Select AI configurations and trace changes from initiation through approval, testing, implementation, and post-change validation. Verify version history and rollback capability exist. | Configuration repository; version history; change tickets; approvals; test results; deployment logs; rollback plans; baseline configurations. |
| 13 | AI Model Change Management | ITGC change management | Ensure model changes, retraining, model-provider updates, fine-tuning, embeddings, indexes, and material AI components are authorized, tested, and approved before production use. | Sample model changes and provider updates. Verify documented impact assessment, testing against defined criteria, independent validation, approval, deployment authorization, and post-change performance review. | Model versions; release records; testing evidence; validation results; approval logs; deployment records; post-implementation testing. |
| 14 | AI Bill of Materials / Traceability | ITGC / detective control | Maintain sufficient traceability of AI components to reconstruct how an output was generated and support investigation, audit, and control testing. | Determine whether selected AI processes retain prompts, outputs, model versions, parameters, system messages, retrieval sources, and plugins. Attempt to trace sampled outputs to the configuration that produced them. | AI bill of materials; prompt/output logs; model version; parameters; plugins; retrieval sources; system messages; configuration history. |
| 15 | AI Deployment Approval | Application / change control | Ensure AI functionality affecting ICFR is independently validated and formally approved before production deployment or material expansion of use. | Select new AI capabilities and trace deployment to approved requirements, test results, independent validation, control-owner approval, and production authorization. Verify unresolved exceptions were appropriately addressed. | Deployment checklist; test scripts; validation results; approval evidence; exception logs; production release records. |
| 16 | Human-in-the-Loop Review | Application control / manual review | Require human corroboration of AI outputs proportionate to financial reporting risk, treating AI-generated outputs as assertions requiring supporting evidence rather than authoritative facts. | Identify ICFR controls using AI outputs. Determine whether reviewers perform documented validation, re-performance, or risk-based sampling and whether review depth reflects financial reporting impact. | Reviewer sign-off; re-performance evidence; sampled transactions; supporting documentation; review checklist; exception resolution. |
| 17 | AI Output Accuracy & Completeness | Application control / validation control | Ensure AI-generated outputs used in ICFR are sufficiently accurate and complete for their intended purpose and within defined tolerances. | Obtain management's performance criteria and independently test a sample of AI outputs against reliable source information. Evaluate accuracy, completeness, exceptions, and whether errors exceeded defined tolerances. | Test populations; source-to-output comparisons; accuracy metrics; completeness testing; exception reports; management review; defined tolerance thresholds. |
| 18 | Hallucination & Source Validation | Application preventive/detective control | Prevent unsupported AI-generated information from influencing financial reporting by requiring source validation, citations, confidence thresholds, or human review. | Sample material AI outputs and inspect whether supporting sources are captured and validated. Test outputs containing unsupported statements and verify escalation or blocking mechanisms operate as designed. | Source citations; retrieval records; confidence scores; validation checklists; reviewer evidence; blocked outputs; exception logs. |
| 19 | AI-Assisted Reconciliation / Matching | Automated application control with HITL | Ensure AI-assisted reconciliations or matching activities only automatically process items within validated thresholds, while exceptions receive appropriate human review. | Reperform the reconciliation logic using test items above, below, and around the approved confidence threshold. Inspect exception routing, reviewer actions, threshold approvals, and post-change testing. | Threshold configuration; test results; matched/unmatched populations; exception queues; reviewer approvals; change approvals; post-change sampling. |
| 20 | AI Data Ingestion & Extraction | Application control | Ensure AI-generated data extraction or ingestion is accurate, complete, authorized, and based on approved source data before information enters an ICFR process. | Test source documents through the extraction process, including standard and edge-case formats. Compare extracted information with source records and investigate low-confidence or incorrect results. | Source documents; extracted data; confidence scores; exception queues; dual-review evidence; accuracy testing; approved source inventory. |
| 21 | AI Transformation & Integration | Application / interface control | Prevent AI-driven transformation, enrichment, mapping, or integration errors from corrupting financial data or downstream reporting processes. | Trace selected transformed records from source through AI processing to downstream systems. Test mapping rules, edge cases, completeness, rejected records, and downstream reconciliation. | Source/target records; mapping rules; transformation logs; interface reports; rejected-item reports; reconciliations; test results. |
| 22 | AI Agents & Automated Actions | Application preventive control | Ensure autonomous AI agents cannot initiate unauthorized financial actions, exceed approved authority, or bypass established approval and authorization controls. | Inspect agent permissions, action boundaries, approval gates, interfaces, and transaction limits. Execute or inspect controlled test scenarios designed to determine whether unauthorized actions are blocked. | Agent configuration; authorization matrix; transaction limits; approval queues; execution logs; blocked-action evidence; test scenarios. |
| 23 | AI Prompt Injection & Manipulation | IT/security/application control | Protect AI-enabled ICFR processes against prompt injection, malicious instructions, manipulated source content, and other inputs capable of altering authorized processing. | Review security and application testing for prompt injection and manipulated inputs. Inspect whether high-risk scenarios are tested and whether unauthorized instructions are prevented from influencing financial outputs or actions. | Security testing; adversarial test cases; prompt filtering rules; incident logs; blocked attempts; remediation evidence; penetration-test results. |
| 24 | AI Data Confidentiality & Leakage | Preventive IT/application control | Prevent unauthorized disclosure or processing of confidential, personal, regulated, or financially sensitive information through GenAI systems. | Identify sensitive data permitted in AI workflows. Test input restrictions, masking, access controls, retention settings, vendor terms, and monitoring for unauthorized data submission or disclosure. | Data classification; DLP rules; masking configurations; access reports; vendor agreements; monitoring logs; incident reports. |
| 25 | AI Output Reliance Determination | SOX management control | Determine whether management relies on AI output as evidence supporting an ICFR control and apply evidentiary requirements proportionate to that reliance. | For each AI-enabled key control, interview the control owner and inspect the actual control workflow. Determine whether AI output is independently corroborated or forms the basis for the control conclusion. | Control narratives; walkthrough documentation; reviewer procedures; AI outputs; re-performance evidence; reliance assessment; management representation. |
| 26 | Evidence Retention for AI-Dependent Controls | SOX evidence / documentation control | Retain sufficient evidence to demonstrate what AI control activity occurred, using which model and configuration, who reviewed it, and how exceptions were resolved. | Select operating instances of AI-dependent controls and verify retained evidence supports occurrence, accuracy, reviewer involvement, timing, configuration, exceptions, and resolution. | Control evidence; prompts; outputs; model versions; configuration records; timestamps; reviewer sign-offs; exception resolution. |
| 27 | Sampling of AI Outputs | SOX application monitoring control | Ensure samples used to validate AI-generated results are appropriately designed, documented, representative, and responsive to the risks of AI variability. | Inspect sampling methodology for AI-dependent controls. Evaluate population completeness, sample rationale, risk-based selection, exceptions, and whether sample size or frequency changes when AI performance deteriorates. | Population reports; sampling methodology; sample selections; rationale; exceptions; reviewer sign-offs; revised sampling decisions. |
| 28 | AI Knowledge Retrieval / Summarization | Application control / review control | Ensure AI-generated summaries or retrieved information used in financial reporting are complete, sourced, current, and appropriately reviewed before reliance. | Test selected summaries against the underlying source population. Evaluate completeness of retrieval, source currency, interpretation accuracy, material omissions, citations, and reviewer challenge of contradictory information. | Source repository; retrieval logs; citations; generated summaries; completeness testing; reviewer comments; contrary-information evidence. |
| 29 | Forecasting & Judgmental AI | Management review control | Ensure AI-supported forecasts and judgments are appropriately challenged, supported by reliable information, and evaluated against actual results before continued reliance. | Inspect forecasts used in ICFR and compare assumptions, source data, AI outputs, reviewer challenge, contrary information, and subsequent actual results. Investigate significant unexplained variances. | Forecasts; assumptions; source data; AI outputs; reviewer challenge; variance analysis; hindsight reviews; management approval. |
| 30 | AI Performance Monitoring | Monitoring control | Continuously monitor AI accuracy, completeness, precision, recall, exceptions, latency, fairness, drift, and other risk indicators against approved tolerances. | Inspect monitoring dashboards and reports for selected AI systems. Verify defined thresholds, monitoring frequency, responsible reviewers, investigated exceptions, and escalation when metrics exceed tolerance. | KPI/KRI dashboards; monitoring reports; threshold definitions; alerts; investigation records; reviewer sign-offs; escalation evidence. |
| 31 | Model Drift Monitoring | Monitoring / detective control | Detect deterioration in AI performance caused by changing data, source formats, business conditions, model behavior, or other environmental changes. | Review drift metrics and trend reports. Compare model performance over time and inspect whether predefined thresholds trigger investigation, retraining, rollback, or additional human review. | Drift reports; historical performance metrics; threshold alerts; trend analysis; retraining decisions; rollback evidence. |
| 32 | Ongoing AI Control Evaluation | SOX monitoring control | Periodically evaluate whether AI-related controls remain properly designed and operating effectively despite changes in models, data, configurations, vendors, and processes. | Perform or inspect management's periodic evaluation of AI-enabled controls. Review historical and hypothetical test cases, adverse scenarios, control-operator performance, and changes since the prior evaluation. | Periodic evaluation reports; test populations; hypothetical scenarios; operator assessments; independent challenge; conclusions; management sign-off. |
| 33 | Independent AI Validation | Monitoring / independent control | Obtain independent challenge or validation for high-risk AI systems where model complexity, financial impact, or reliance warrants additional assurance. | Identify high-risk AI use cases and determine whether independent validation is required. Inspect validation scope, methodology, findings, management response, and closure of identified issues. | Independent validation report; testing methodology; findings; management responses; remediation evidence; validation approval. |
| 34 | AI Exception Management | Application monitoring control | Ensure AI-generated exceptions are identified, investigated, resolved, and appropriately escalated rather than automatically accepted or ignored. | Sample AI-generated exceptions and trace them to investigation, resolution, reviewer approval, and closure. Evaluate whether unresolved exceptions remain visible and are appropriately escalated. | Exception reports; investigation records; resolution evidence; reviewer sign-offs; aging reports; escalation records. |
| 35 | AI Control Deficiency Assessment | SOX deficiency evaluation control | Ensure AI-related control failures are evaluated for severity, root cause, scope, and financial reporting impact using SOX deficiency assessment criteria. | Review identified AI incidents and control failures. Determine whether management evaluated severity, affected populations, financial reporting implications, root causes, compensating controls, and potential broader impact. | Deficiency assessments; root-cause analysis; affected-population analysis; SOX evaluation; management conclusions; legal/compliance consultation. |
| 36 | AI Remediation & Retesting | SOX remediation control | Ensure identified AI control deficiencies are remediated timely and retested to demonstrate that corrective actions operate effectively. | Select AI-related deficiencies and trace remediation from root cause through corrective action, implementation, retesting, and closure. Verify remediation addresses the underlying failure rather than only the symptom. | Remediation plans; corrective-action records; revised configurations; retesting results; owner sign-off; closure approval. |
| 37 | AI Incident Escalation | Governance / monitoring control | Ensure material AI incidents, control failures, configuration changes, privacy events, or reliability issues are promptly communicated to appropriate management and governance bodies. | Inspect incident procedures and sample material AI incidents. Verify escalation occurred within defined timelines and included sufficient information regarding impact, root cause, remediation, and residual risk. | Incident tickets; escalation matrix; notification records; governance minutes; impact assessments; remediation plans. |
| 38 | Third-Party GenAI Providers | Vendor management / ITGC control | Ensure third-party AI providers supporting ICFR meet defined security, reliability, availability, data, change-management, and assurance requirements. | Identify third-party AI providers used in SOX processes. Review contracts, SOC reports or equivalent assurance, service changes, data provisions, incidents, and management's assessment of complementary controls. | Vendor inventory; contracts; SOC reports; assurance reports; security assessments; vendor-change notices; incident records. |
| 39 | Third-Party Model Changes | Vendor change management | Detect and assess provider-driven model or service changes that could alter AI performance, control behavior, or financial reporting risk. | Review vendor change notifications for selected providers. Determine whether changes were assessed for impact, independently tested where necessary, approved, and monitored after implementation. | Vendor notices; impact assessments; test results; approvals; updated risk assessments; post-change monitoring. |
| 40 | External Communication & Disclosure | Management review / disclosure control | Ensure material GenAI impacts, limitations, incidents, and risks communicated externally are accurate, complete, appropriately approved, and consistent with applicable requirements. | Identify financial reporting or external disclosures materially affected by GenAI. Inspect supporting evidence, management review, approval, source validation, and consistency with known AI limitations or incidents. | Disclosure drafts; supporting analysis; AI-use documentation; management review; legal/compliance approval; final disclosure evidence. |
Recommended SOX audit hierarchy
From a practical SOX perspective, I would not test all 40 controls with equal intensity. I would organize them into the following control hierarchy.
Tier 1 Entity-level controls
These are the controls I would assess first because deficiencies here can affect multiple AI-enabled processes:
- GenAI Governance & Accountability
- GenAI Acceptable Use & Ethical Boundaries
- GenAI Risk-Based Scoping
- GenAI Risk Assessment
- AI Fraud Risk
- AI Governance Competence
- AI Performance Accountability
The source emphasizes that GenAI's accessibility and ability to bypass traditional approval channels make the control environment particularly important. It also identifies ownership, authority, escalation paths, and defined scope as common controls across AI capabilities.
Tier 2 GenAI ITGCs
These should generally be incorporated into the existing SOX ITGC framework rather than creating an entirely separate "AI audit":
- AI Access Management
- Segregation of Duties
- AI Configuration Management
- AI Model Change Management
- AI Bill of Materials / Traceability
- AI Deployment Approval
- Third-Party GenAI Providers
- Third-Party Model Changes
The key audit principle is: if the AI component can change the outcome of an ICFR control, treat the AI component as part of the relevant technology environment and assess access, change, operations, and evidence controls accordingly.
The source expressly states that GenAI models, configurations, fine-tuning artifacts, embeddings and RAG indexes should be treated as configuration items subject to access control, segregation of duties and change management.
Tier 3 AI-enabled application controls affecting ICFR
These are the controls where I would expect the largest change to traditional SOX testing:
- AI Deployment Approval
- Human-in-the-Loop Review
- AI Output Accuracy & Completeness
- Hallucination & Source Validation
- AI-Assisted Reconciliation / Matching
- AI Data Ingestion & Extraction
- AI Transformation & Integration
- AI Agents & Automated Actions
- AI Prompt Injection & Manipulation
- AI Data Confidentiality & Leakage
- AI Output Reliance Determination
- Evidence Retention
- Sampling of AI Outputs
- Knowledge Retrieval / Summarization
- Forecasting & Judgmental AI
The source makes an important distinction for SOX purposes: AI assistance is not necessarily AI reliance. If a control owner independently re-performs or validates the AI result, reliance may be limited. If management accepts the AI output as the evidence supporting the control conclusion, the AI becomes part of the control's evidentiary foundation.
That distinction should be explicitly incorporated into the SOX walkthrough.
Tier 4 Monitoring and deficiency manageme
- AI Performance Monitoring
- Model Drift Monitoring
- Ongoing AI Control Evaluation
- Independent AI Validation
- AI Exception Management
- AI Control Deficiency Assessment
- AI Remediation & Retesting
- AI Incident Escalation
This is particularly important because a traditional ITGC conclusion of "change management is effective" may not be enough for GenAI. A model can remain technically unchanged while its behavior changes because the underlying data, retrieval corpus, prompts, vendor model, external environment, or business conditions change. The source therefore emphasizes continuous monitoring and separate evaluations.
Is Your Team Ready? Six Questions for Your Next Steering Meeting
Run a short readiness test and answer each question with evidence and not with confidence. Can you identify every AI use that affects financial reporting? Can you explain each AI-enabled control in plain language to someone outside the team? Can you reproduce a prior result? Can you show meaningful human review? Can you demonstrate controlled change? Can you prove what happens when the AI fails? A no on any of them points to the section of this article that addresses it.
Add a measurement you can compute yourself. This is a reader exercise, and the numbers are yours to fill in. Divide the number of SOX-relevant AI uses that have a named owner, a mapped assertion, and a current evidence standard by the total number of AI uses you found in finance processes, and record the result with the date and the data source. Run it each quarter and watch the trend. A low first number is normal, because the first sweep usually finds uses nobody expected.
Turn the answers into a ninety day plan with owners. Fix inventory gaps and unowned uses in the first thirty days, because everything else depends on them. Rewrite control descriptions and set the evidence standard in the next thirty, and use the last thirty for validation, monitoring thresholds, and the failure tabletop. Report progress to the audit committee in the same quarter, and tell the external auditor what you found and when you will finish. Early disclosure of a known gap with a dated plan is a stronger position than a gap the auditor discovers.
The reproduction test is more revealing than any ratio. Pick one result from the prior quarter and ask the team to rebuild it from retained evidence in one working day. Time the exercise and note every missing artifact, then give each gap an owner. If the team cannot do it, an auditor will find that out within an hour of fieldwork, and it is far better to learn it in your own conference room.
Final Perspective
The next decision on your desk is probably a supplier update, a new finance assistant, or a request to let an agent post routine entries. Each of them should now pass four checks before it goes live: an inventory entry with an owner and an assertion, a control description in operational language, a change record that names who can alter the behavior, and a reviewer who will see the underlying evidence. Auditors will keep asking the five questions in different words, and the organizations that answer them well will have turned them into ordinary control mechanics. COSO's framing helps here, since it places AI inside the control structure you already run.
The takeaway is to schedule the one-day reproduction test for next month and make a person accountable for the result. Pick the AI-enabled control with the highest financial reporting impact and try to rebuild a prior result from retained evidence alone. What you cannot rebuild is your real control gap, and it will cost far less to find it now than during fieldwork.
References
- COSO, Achieving Effective Internal Control Over Generative AI (GenAI), https://www.coso.org/generative-ai
- Deloitte, Heads Up, COSO Releases Publication on Internal Controls Related to Generative AI, https://dart.deloitte.com/USDART/home/publications/deloitte/heads-up/2026/coso-internal-controls-generative-ai
- PCAOB, Spotlight: Staff Update on Outreach Activities Related to the Integration of Generative Artificial Intelligence in Audits and Financial Reporting, https://pcaobus.org/documents/generative-ai-spotlight.pdf
- EY, 2025 AICPA & CIMA Conference on Current SEC and PCAOB Developments, Compendium of significant accounting, auditing and reporting issues, https://www.ey.com/content/dam/ey-unified-site/ey-com/en-us/technical/accountinglink/documents/ey-sec29318-251us-12-13-2025.pdf
- Grant Thornton, The unexpected obstacle to AI-enabled SOX and internal audit, https://www.grantthornton.com/insights/articles/advisory/2026/the-unexpected-obstacle-to-ai-enabled-sox-and-internal-audit
