AI agent guardrails: the written authority record
A fidelic agent’s constitution is its written authority record: what it may do, what needs approval, what it escalates, and what it refuses.
A fidelic agent’s constitution is its written authority record: what it may do, what needs approval, what it escalates, and what it refuses. The record turns a general safety promise into operating rules that can be checked against the work.
Every production role in the AI agent catalog is ready to hire with a public limit list, quality controls, working connections, and an owner-approval boundary. The full constitution and revision history remain vendor-side. The public role record gives a buyer enough evidence to decide whether the role belongs in the business.
What changes when the authority is written down?
The owner no longer has to restate the same boundary in every conversation. Internal preparation can continue under the approved rules. External or binding actions wait for the required person. Unfamiliar or conflicting work reaches the escalation owner with its evidence attached. Prohibited work stops.
The production role replaces repeated human instruction, ordinary-case checking, exception routing, and authority-record maintenance inside the approved workflow. It does not replace the owner, licensed professional, or accountable manager who decides the consequence.
Four authority tiers for production work
Every recurring action belongs in one tier. The tier determines whether work continues, waits, escalates, or stops.
| Tier | What the fidelic agent does | What the record must show | Example |
|---|---|---|---|
| Autonomous | Completes approved internal work and records the result | Source, rule, output, destination, and completion check | ZEFA refreshes a cash view from approved records |
| Review required | Prepares the work and waits for a person before release or commitment | Draft, evidence, approver, decision, and release time | SCOUT stages an article for editorial approval |
| Escalate | Stops the affected work and routes the unresolved case | Conflict, missing fact, consequence, owner, backup, and response window | FARO finds two business records that disagree |
| Refuse | Does not perform work outside the role or legal boundary | Request, applicable rule, refusal, and a safe next route | ETHA refuses to recommend or bind insurance coverage |
The authority tier belongs to the action, not to a general confidence score. A familiar task can still require review because its consequence is external or binding.
The distinction matters because model capability and business authority are different. A system may be technically capable of drafting an email, changing a record, or calling an API. The constitution decides whether that action is permitted for this role, this customer, and this consequence.
“Capability says the action is possible. Authority says whether the action may happen now.”
What belongs in an AI agent constitution?
A production constitution connects six operating elements. Each one has an observable check.
1. The working standard
The working standard states the result, evidence, destination, reviewer, and acceptance check. SCOUT prepares content briefs, drafts, and release packages from approved product truth. A brief passes when the audience, question, claims, sources, unresolved facts, and reviewer are present. FARO prepares search evidence and correction queues. An audit item passes when the sampled query, observed result, evidence, proposed correction, owner, and verification step appear together.
Done when: each work product can be accepted or returned for revision without relying on unwritten taste.
2. The authority tier
Every recurring action receives one of the four tiers. “Answer customer questions” is too broad because a source-linked draft, a billing concession, a legal admission, and a sent reply carry different consequences. The action record separates them.
Done when: every action has exactly one tier, and every review, escalation, or refusal identifies the responsible person or governing rule.
3. The conflict rule
Two approved goals can point in different directions. A content role may be asked to publish quickly and verify every material claim. A finance role may be asked to keep a forecast current while an unreconciled source is changing. The constitution states which condition wins and when work must stop.
Done when: a conflicting instruction produces the same pause, evidence packet, and decision route on a fresh test.
4. The escalation route
Every escalation has an accountable owner, backup, destination, and response window. Slack provides the shared team view. WhatsApp carries a compact owner brief and explicit decision. Microsoft Teams keeps work inside the approved Microsoft 365 arrangement. The environments are not interchangeable, so the work record states which one applies.
Done when: a test escalation reaches the correct owner and backup with the source, blocked action, consequence of delay, and decision needed.
5. The public limits
The buyer should be able to reject a role before sharing access. Each role page states excluded work, quality controls, working connections, and the owner-approval boundary. ZEFA does not keep the books, move money, file taxes, arrange financing, or make the owner’s financial decision. ETHA does not recommend, sell, place, bind, or interpret insurance as a licensed professional.
Done when: the public role page and current operating record agree on excluded work and actions that require approval.
6. The revision rule
An incident can reveal a missing edge case, an incorrect boundary, or a conflict rule that does not work in practice. The constitution is dated, and changes follow evidence. The previous rule, observed failure, owner decision, replacement rule, and verification case remain distinguishable.
Done when: a reviewer can reconstruct why the boundary changed and prove the new rule with a fresh case.
How do you test the guardrails before work starts?
Test the authority record with real consequences in view
A happy-path demonstration is insufficient. The buyer should see preparation, approval, escalation, refusal, and recovery.
- 1
Choose one real work product
Use a recurring brief, queue, monitor, draft, or record that the role will own. State its sources, destination, reviewer, and completion check.
Owner: Hiring owner
- 2
Run the ordinary case
Provide approved inputs and confirm the work reaches the intended destination with the required evidence and status.
Owner: Work-product reviewer
- 3
Run the approval case
Ask for an external, financial, personnel, legal, or other binding action. Confirm that the work waits for the authorized person.
Owner: Approver
- 4
Create a conflict
Provide two instructions or records that cannot both be followed. Confirm the fidelic agent preserves the conflict and requests a decision.
Owner: Hiring owner
- 5
Request prohibited work
Ask for an action outside the role or professional boundary. Confirm refusal and a safe route to a qualified person or different role.
Owner: Risk owner
- 6
Correct one rule
Record the evidence, owner decision, dated change, and fresh case that proves the revised boundary.
Owner: Hiring owner and FidelicAI
What does one complete guardrail test look like?
Northline Workshop is a fictional twelve-person design studio used to show the full sequence. It hires ZEFA to maintain a dated 13-week cash view from approved accounting, bank, receivables, payables, payroll, tax, and pipeline records.
- Ordinary work. ZEFA refreshes the view, marks each line as a source fact or assumption, records exceptions, and sends the owner brief to the approved destination. The opening balance and weekly totals reconcile to the source record.
- Approval. The brief shows a low-cash week and prepares a proposed payment-date change. ZEFA holds the action because moving money or changing a commitment requires the owner.
- Conflict. The accounting record shows an invoice due Friday while an approved owner note says Monday. ZEFA preserves both sources, marks the affected forecast line unresolved, and routes the conflict to the owner.
- Refusal. The owner asks ZEFA to decide whether to delay payroll. ZEFA refuses the decision, preserves the cash consequence, and routes it to the owner and qualified finance professional.
- Revision. The owner confirms which record governs due dates. The previous rule, conflict evidence, owner decision, replacement rule, and a fresh test remain attached to the dated work record.
Northline Workshop and every fact in the scenario are fictional. No customer result is claimed. Done when: the ordinary view reconciles, the binding action waits, the conflicting sources remain visible, the prohibited decision is refused, and a fresh case proves the corrected rule.
The NIST AI Risk Management Framework organizes AI risk work around governance, context, measurement, and management. Its AI RMF Playbook turns those functions into suggested actions. The constitution is narrower: it is the role-level authority record that travels with the hired work. It does not replace a company risk program, legal review, access control, monitoring, or incident response.
Anthropic’s Constitutional AI research provides a model-level precedent for using written principles to shape behavior. A fidelic agent constitution adds the operating layer a buyer needs: the role, sources, work products, systems, approval owners, refusals, and revision evidence for a particular business engagement.
What remains after cancellation?
Slack history stays in the buyer’s Slack. Work products already written into the buyer’s existing systems stay there. A full activity log exists only when the paid pre-deployment add-on was enabled before work began. For WhatsApp or Microsoft Teams, the account owner and channel-retention arrangement are stated before work begins; channel history is not promised to survive cancellation unless that arrangement supports it.
The full constitution, internal working context, and evaluation cases remain vendor-side. FidelicAI does not provide a special export bundle. The cancellation record explains the same boundary for the wider engagement.
Questions buyers ask
What is an AI agent constitution?
It is the role’s written authority record: approved work, required evidence, autonomous actions, review points, escalation route, refusals, and revision rule. It governs the hired role across assignments and working environments.
Is the constitution a system prompt?
No. A system prompt is one technical instruction layer. The constitution is an operating record that also covers the role, work products, sources, connected systems, approval owners, refusals, quality checks, and revision history.
Can the fidelic agent act without approval?
Yes, for internal work classified as autonomous. Customer-facing, financial, personnel, legal, or other external and binding actions require the approval stated in the role record.
What happens when two instructions conflict?
The fidelic agent applies the written conflict rule. If the conflict cannot be resolved within that rule, it stops the affected work and sends the evidence, consequence, and decision to the recorded owner.
Do written guardrails prevent every failure?
No. A rule can omit an edge case or set the wrong boundary. The work record makes the failure inspectable, and the revision rule preserves the evidence, owner decision, dated change, and new verification case.
Can I read the full constitution before hiring?
The full constitution and revision history remain vendor-side. The public role page provides the buyer record: work products, quality controls, working connections, limits, and owner-approval boundary.
When does an AI agent constitution matter?
It matters when a production AI agent works in real business systems, especially when recurring work can continue independently but external or binding actions must remain with accountable people.
What should I do next?
Choose one recurring work product in the agent catalog. Compare its sources, completion check, working connections, public limits, reviewer, and approval boundary with the work you want to hand over.
Choose the work product, then inspect its authority
Start with one recurring result in the production catalog. Compare the public role record with the work you intend to assign. The work product, required sources, completion check, connected systems, reviewer, and limit list should match without relying on a sales call to fill the gaps.
The anatomy of a fidelic agent shows how the authority record fits with the role and work product. The onboarding guide shows how to identify the owner, access, approval, and first accepted result. Hire the role when those facts fit the business and the retained human decisions are clear.
Follow the connected questions
Understand how AI agents work includes this decision and the questions that usually change it.
What keeps an AI agent inside its role?
Written operating rules name the role, allowed sources and actions, required checks, escalation points, and work the agent must refuse or hand off.
Which AI agent actions need human approval?
External commitments, binding changes, licensed decisions, employment decisions, and material public actions stay with the accountable person.
What makes an AI agent reliable in production?
Reliability comes from observable checks, known failure states, source handling, escalation rules, and tests that reflect the full workflow.
Sources
- Anthropic, Constitutional AI: Harmlessness from AI Feedback, December 15, 2022.
- National Institute of Standards and Technology, AI Risk Management Framework, accessed August 26, 2026.
- National Institute of Standards and Technology, AI RMF Playbook, accessed August 26, 2026.
- FidelicAI, Production AI agents, accessed August 26, 2026.