The fastest way to get burned: “turn it on for everything”
If you’ve ever watched an eager stakeholder route around procurement because “the tool can just buy it,” you already know the failure mode: autonomy gets treated like a feature toggle. Then the first ugly edge case hits—wrong entity, wrong tax treatment, a supplier that shouldn’t have been used, or a renewal that quietly changed liability—and suddenly everyone rediscovers why delegation of authority (DoA) exists.
The governance move that actually works is boring but effective: autonomy is granted per work type, per category, per value band, per supplier tier. And it’s earned through evidence (clean outcomes, stable pricing, predictable specs, low dispute rates), not asserted because a vendor demo looked smooth.
Think in ladders, not switches: Recommend → Execute → Autonomous
Procurement tasks don’t fail in the same way. Buying a known SKU from an approved supplier is nothing like awarding a new contract. So give AI a ladder to climb. Each rung has a different control posture, different audit expectations, and different “blast radius” when something goes wrong.
Rung 1 — Recommend (AI suggests, humans decide)
Use this rung when the decision is consequential, ambiguous, or politically sensitive. AI can draft the comparison, flag risk, and propose an action—but a human approves the supplier, the commercial position, and the final document.
Good fits: sourcing plan suggestions, RFx draft questions, supplier shortlists for known categories, contract clause redlines for review, renewal risk summaries, spend classification clean-up.
Controls that matter: show-your-work citations to internal policy and contract repository; reason codes for recommendations; conflict-of-interest checks; a mandatory human sign-off captured in the workflow.
What you learn here: whether the AI’s suggestions are consistently sensible in your environment (your templates, your risk rules, your suppliers), not in a lab.
Rung 2 — Execute (AI performs steps inside a fenced process)
Execution autonomy is where teams get overconfident. The trick is to allow AI to do the mechanical work while keeping the decision points gated. “Execute” means the agent can create requisitions, populate POs, chase confirmations, and route approvals—only within pre-approved guardrails.
Good fits: reorder of catalog items; PO creation against an existing contract; scheduling deliveries; three-way match exception triage with predefined playbooks; chasing overdue ASN/invoices; creating change requests for human approval.
Guardrails: only approved suppliers; only contracted items/rate cards; only within a defined budget and cost center; only using pre-approved templates; hard stop when data is missing or contradictory.
Audit needs: a complete activity log (who/what/when), plus the input data used (contract version, price file, tax rules, ship-to) so Finance and Internal Audit can reconstruct the decision path.
Rung 3 — Autonomous (AI decides and executes within a small box)
True autonomy is acceptable only where outcomes are predictable and the downside is capped. That usually means low-value, low-risk, highly standardized buys with stable suppliers and clean master data. If you can’t explain the control set in one breath, it’s not ready.
Good fits: office consumables, standard IT peripherals from an existing framework, routine facilities supplies, pre-approved training seats from a contracted provider, low-risk MRO where spec is unambiguous.
Non-negotiables: automatic budget checks; duplicate detection; price variance thresholds; segregation of duties (the agent can’t both create and approve); and a kill switch that actually works.
Operating model: autonomy is time-bound (e.g., 90 days), reviewed, and renewed only if the evidence stays clean.
Autonomy is earned per category, per value band, per supplier tier
The same spend amount can be “fine” in one category and reckless in another. A $5,000 PO for toner is noise; $5,000 for a niche consulting engagement can quietly create IP, data access, and misclassification risk. So define autonomy in a matrix, not a statement.
A practical matrix you can actually run
Start with three supplier tiers and three value bands, then map the ladder rung allowed per cell. Keep it simple enough that stakeholders can remember it.
Supplier tier: Tier 1 (strategic/critical), Tier 2 (approved/preferred), Tier 3 (not approved / new / one-off).
Value bands: Band A (≤ $1,000), Band B ($1,001–$10,000), Band C (> $10,000). Adjust to your DoA, but don’t pretend one number fits all categories.
Category risk overlay: mark categories as Low/Medium/High based on regulation, safety, data access, IP, and service complexity. High-risk categories drop one rung automatically.
Example policy that doesn’t collapse under pressure: Tier 2 + Band A + Low-risk category can be Autonomous. Tier 2 + Band B is Execute (human approval at the commit point). Anything Tier 1 is Recommend at best unless it’s a tightly defined call-off under an existing agreement. Tier 3 is Recommend only—because “new and unknown” is where automation creates the biggest mess.
Hard stops: what should never be autonomous
Some decisions carry legal and ethical weight that you do not want delegated to a system that can’t be deposed, can’t be cross-examined, and can’t be held accountable in the way an employee can. Even if the AI is “right” most of the time, the one time it isn’t can be catastrophic.
New supplier onboarding: KYC/AML checks, sanctions screening, beneficial ownership, bank detail verification, data processing agreements—this stays human-owned, with AI assisting on document collection and completeness checks.
Contract award: the final selection and award decision requires a human accountable owner, documented rationale, and conflict-of-interest attestation.
Price negotiation and commercial concessions: AI can propose positions and draft emails, but a human negotiator owns the strategy and makes commitments.
Anything above the delegation threshold: if your DoA says a person must approve, an agent cannot be the approver. At most it can prepare the pack and route it.
Policy exceptions: if the buy requires an exception (single-source justification, emergency purchase), autonomy stops and the workflow escalates.
Evidence gates: what you must prove before climbing a rung
Teams often argue autonomy from intuition: “It’s low risk.” That’s how you end up with a tool buying the right thing the wrong way. Treat each rung increase like a controlled release. You need proof that the process, data, and suppliers behave predictably.
Master data health: supplier status, bank details, tax codes, ship-to locations, and contract metadata are clean and current. If your vendor master is messy, autonomy will amplify the mess.
Price stability: a defined variance rule (e.g., stop if price differs from contract by more than X% or exceeds a not-to-exceed amount). Pick X per category; don’t set one global number.
Specification clarity: the item/service can be described unambiguously. If stakeholders routinely change scope after ordering, you’re not ready for autonomy.
Dispute history: low rate of invoice disputes, returns, and delivery exceptions for the supplier/category combination.
Control performance: the agent reliably triggers approvals, flags exceptions, and produces an audit trail that Finance accepts without detective work.
Two concrete scenarios (and where autonomy breaks)
Scenario A: Contracted laptops for new hires
This is where autonomy can shine. You have a framework agreement, a fixed configuration list, and an approved supplier. Let the agent execute: create the PO, confirm delivery dates, and reconcile invoices. Put a hard stop on anything outside the contracted SKUs or above the per-unit ceiling price. If the hiring manager requests a “slightly better model,” autonomy stops—because that’s a spec change with commercial implications.
Scenario B: Marketing asks for a “quick” freelance designer
This looks small-dollar and urgent, which is exactly why it’s dangerous. The real risk is IP ownership, data access, and worker classification. Even at $800, this should be Recommend: AI can suggest pre-approved agencies, draft a scope, and assemble the onboarding pack. A human decides whether it’s a contractor, an agency engagement, or something that needs Legal.
The mistake people repeat: autonomy without ownership
Autonomy programs fail when nobody is clearly on the hook for outcomes. If an agent places an order that violates policy, “the system did it” is not an acceptable post-mortem. Assign a process owner per category who can pause autonomy, adjust thresholds, and explain the control design to Audit without hand-waving.
If you want a simple rule that holds up: let AI do high-volume work where the decision is already made (contracted, approved, low-risk). Keep humans on anything that creates a new obligation, a new supplier relationship, or a new commercial position. That’s not anti-AI. That’s how you keep the benefits without inviting the kind of failure that gets autonomy banned for everyone.