Your risk management framework probably wasn't designed for this question. But the distinction matters more than you think.
When a manager uses an AI agent to draft a vendor risk assessment, analyze control test results, or recommend audit priorities, you're facing a choice: treat that AI output as reference material requiring human judgment, or accept it as a delegated decision. The path you choose determines what controls you need, how you test them, and whether your oversight model actually works.
The Decision You're Facing
You need to classify how AI participates in your organization's decision-making. This isn't about whether AI is "good" or "bad." It's about control design. The wrong classification creates gaps in your assurance framework.
Three factors drive this choice: the decision's materiality, the AI system's transparency, and your team's ability to validate outputs. Get the classification wrong, and you'll either over-control low-risk automation or under-control material judgments that look automated but aren't.
Key Factors That Affect Your Choice
Decision materiality: Does the output directly affect financial reporting, regulatory filings, or control effectiveness conclusions? If yes, you're designing controls for a material decision process, not a productivity tool.
System transparency: Can you trace how the AI reached its conclusion? Some systems provide decision logic; others operate as black boxes. ISO 31000's emphasis on understanding decision processes applies here. If you can't explain the reasoning chain, you can't assess the decision quality.
Validation capacity: Does your team have the expertise and time to verify AI outputs before acting on them? This isn't about spot-checking. It's about whether substantive review is feasible given your resources and the decision's time sensitivity.
Information reliability: Where does the AI get its data? If it's pulling from your GRC platform's control testing results, that's different from scraping public sources or using training data you can't audit.
Path A: Treat AI as Reference Material
Choose this path when you need human judgment to remain the decision point.
When this applies:
- The decision affects regulatory obligations, audit findings, or control effectiveness ratings
- You can't verify the AI's data sources or reasoning process
- The output requires professional skepticism (think: evaluating control design adequacy or assessing third-party risk)
- Your team has the expertise to substantively review and challenge the AI's work
What this means for controls:
You're designing controls around human decision-making that happens to use AI-generated inputs. Your control is the review step, not the AI itself. Document that a qualified person reviewed the AI output, applied independent judgment, and reached their own conclusion.
For example: An AI agent summarizes vendor security questionnaire responses and flags potential gaps. Your third-party risk analyst reviews the summary, examines the original questionnaire, and makes the risk rating decision. The control is the analyst's review and independent judgment, not the AI's flagging logic.
Testing requirements:
Your control testing examines whether reviewers are actually applying judgment. Pull samples where the AI flagged an issue and verify the analyst didn't just accept the flag. Look for instances where the human reached a different conclusion than the AI suggested. If you never see disagreement, your "review" control isn't operating.
Path B: Accept AI as a Delegated Decision
Choose this path when the AI makes the call and humans monitor for exceptions.
When this applies:
- The decision is routine, repeatable, and low-materiality
- You've validated the AI's logic against known-good outcomes
- You have monitoring controls that detect when the AI produces unexpected results
- The decision doesn't require professional skepticism or regulatory judgment
What this means for controls:
You're designing controls around the AI system itself. That means validating the training data, testing the decision logic, and monitoring outputs for drift or bias. Your control environment needs automated control testing that flags anomalies.
For example: An AI agent automatically categorizes policy exceptions by risk level based on defined criteria. It routes high-severity exceptions to the Chief Compliance Officer and standard exceptions to department heads. You've tested the categorization logic, validated it against historical decisions, and set up alerts for unusual patterns.
Testing requirements:
Test the AI's decision accuracy against a validation set. Run known scenarios through the system and verify the outputs match expected results. Set up Key Control Indicators that track decision patterns over time. If the AI suddenly starts routing more exceptions to high-severity, your monitoring control should flag it.
Monitor for bias. If your AI agent consistently rates certain vendor types as higher risk without objective justification, that's a control failure even if the system is technically "working."
Path C: Hybrid Approach with Defined Thresholds
Sometimes you need both paths, with clear triggers that escalate AI decisions to human review.
When this applies:
- You want automation efficiency for routine decisions but human judgment for edge cases
- You can define objective thresholds that separate routine from material decisions
- Your GRC platform can route decisions based on those thresholds
What this means for controls:
Design your control framework with explicit escalation criteria. The AI handles decisions below the threshold; humans review everything above it. Your control is the threshold logic plus the human review process.
For example: An AI agent scores control test failures for severity. Failures below a defined threshold get auto-routed to control owners for remediation. Failures above the threshold or involving entity-level controls get escalated to the Chief Audit Executive for review. Your controls are: (1) the threshold logic, (2) the routing mechanism, and (3) the CAE's review of escalated items.
Testing requirements:
Test both the automation and the escalation. Verify that low-severity items are being handled appropriately through the automated path. Then verify that items meeting escalation criteria actually reach human reviewers and receive substantive attention.
Summary Matrix
| Factor | Path A: Reference Material | Path B: Delegated Decision | Path C: Hybrid |
|---|---|---|---|
| Decision materiality | Material to compliance, audit, or financial reporting | Routine, low-materiality | Mixed; threshold-based |
| Primary control | Human review and judgment | AI system validation and monitoring | Threshold logic + conditional human review |
| Testing focus | Evidence of independent human judgment | AI decision accuracy and bias detection | Both automation and escalation processes |
| Professional skepticism required | Yes | No | Yes, for escalated items |
| Suitable for regulatory obligations | Yes | No | Depends on threshold design |
| Resource intensity | High (human review time) | Medium (setup and monitoring) | Medium to high |
The choice you make determines whether your next audit finds a control gap or a well-designed assurance process. Start by asking decision-makers how they're actually using AI, not how they think they should be using it. That's the insight Grant Purdy would tell you matters most.




