LangGraph vs OpenAI Agents SDK vs CrewAI vs AutoGen: who gives you an approval step

Your engineers want to build the agent on a framework. Fine. Before they pick one, ask a question from your side of the table. Can a person approve an action before it runs? And while they decide, what happens to the run? Four popular frameworks answer differently. Here is what their own docs say.
In one sentence: every framework can pause for a person, but only some pause before the action and survive a restart, and none brings the rules or the record.
Where it comes from
Frameworks started as ways to chain model calls and tools. Human review arrived as agents started changing things. Each framework bolted it on in its own style. Some pause the graph. Some ask a stand-in "user" agent. Some review the output at the end.
These are developer tools, with docs written for developers. This page reads them for the person who signs off on the agent.
What each does, per the docs
LangGraph. "Interrupts allow you to pause graph execution at specific points and wait for external input before continuing"1. When one fires, "LangGraph saves the graph state using its persistence layer and waits indefinitely until you resume execution"1. The saved state lives in a checkpointer, so a paused run can wait for a person across a restart.
One rule matters for money. On resume, "the node restarts from the beginning of the node where the interrupt was called", so "any code before the interrupt runs again"1. If the payment call sits before the interrupt in the same node, it runs twice.
OpenAI Agents SDK. Tools declare when they need approval. "Run results surface pending approvals as interruptions, and RunState lets you serialize paused runs and resume them after decisions are made"2. The approval surface "is run-wide", covering tools reached through handoffs and nested agents2. The SDK also has tool guardrails. Its docs advise using them for "checks before and/or after each custom function-tool call" rather than relying only on agent-level input and output guardrails3.
CrewAI. A task has a human_input flag: "Whether the task should have a human review the final answer of the agent"4. It defaults to off. Tasks also accept guardrails, a "function to validate task output before proceeding to next task"4. Both work on the task's output. Neither is described as a pause before a specific tool call.
AutoGen (AgentChat). Two ways in: during a run through a UserProxyAgent, or between runs5. The docs are frank about the first. A UserProxyAgent "blocks the execution of the team until the user provides feedback or errors out". That puts "the team in an unstable state that cannot be saved or resumed"5. They recommend it only for short interactions, "such as asking for approval or disapproval with a button click"5.
Side by side
| Question | LangGraph | OpenAI Agents SDK | CrewAI | AutoGen AgentChat |
|---|---|---|---|---|
| Pause before a tool call? | Yes, interrupt in the node | Yes, tool needs approval | Not as documented. Reviews task output | Via UserProxyAgent, blocking |
| Paused run survives restart? | Yes, with a checkpointer | Yes, RunState serialises | Not applicable | No, per its docs |
| Resume gotcha | Node re-runs from its start | Keep approval state on the server | Review comes after the work | Team cannot be saved while waiting |
| Who approves? | Whoever your app sends it to | Whoever your app sends it to | Whoever runs the crew | Whoever is at the keyboard |
| Record of the decision | Your code | Your code, plus tracing | Your code | Your code |
What they do well, and where they stop
Do well. LangGraph and the Agents SDK have the right shape for control. The pause sits before the tool call. The run is saved. The person decides. The run resumes with their answer. For a team building one agent, that is solid ground.
Where they stop. The last two rows. A framework gives you a pause and hands the decision to whatever your application does next. It does not name the approver, route to cover, enforce a time limit, sign the rules, or write a record your auditor can verify. Every one of those is code your team writes and maintains.
The per-agent problem. The approval lives in each agent's code. Ten agents, ten implementations. If one team forgets the interrupt on a new payment tool, nothing outside that agent notices.
Output review is not action review. CrewAI's human_input is useful for drafts. For an agent that pays, the review has to come before the tool runs.
What to check
- /01
Where exactly is the pause? Before the tool call, or after the task's output?
- /02
Does a paused run survive a deploy or a crash? Show me.
- /03
On resume, does any code run twice? Where are the side effects?
- /04
Who receives the approval, and how is that person chosen?
- /05
Where is the decision recorded, and can it be edited?
- /06
Does every agent that can write implement the same approval? Who checks?
Where it is going
Frameworks are converging on the LangGraph and Agents SDK pattern: declare which tools need approval, pause, serialise, resume. Our view: that settles the pause and leaves policy open. Teams running several agents will want the approval rules in one place, outside any single framework. That is the same pressure that moved identity out of each application.
Gatehouse fit
Gatehouse sits under whichever framework you pick. The framework runs the loop. Gatehouse checks each action against one set of signed rules before it reaches your systems, sends holds to a named approver, and keeps a record chained by SHA-256. Your engineers keep their framework. You get one approval policy across every agent. If you are building one agent with one approval, a framework may be enough; build or buy walks through the choice.
At a glance
| Category | Vendors and tools |
|---|---|
| Frameworks covered | LangGraph, OpenAI Agents SDK, CrewAI, AutoGen AgentChat |
| Best documented pause-before-action | LangGraph interrupts. Agents SDK tool approvals |
| What none provides | Signed rules, named approver routing, a verifiable record |
| Typical owner | Engineering picks the framework. The process owner owns what needs approval |
| The one test | Pause a run for approval, restart the service, then approve. What happens? |
Sources
- [1]Interrupts, LangGraph docs (LangChain)docs.langchain.com In text
- [2]Human-in-the-loop, OpenAI Agents SDK docsopenai.github.io In text
- [3]Guardrails, OpenAI Agents SDK docsopenai.github.io In text
- [4]Tasks, CrewAI docsdocs.crewai.com In text
- [5]Human-in-the-loop, AutoGen AgentChat user guide (Microsoft)microsoft.github.io In text
Read next

Should you build or buy AI agents?
Building one agent is quick. Running it for years is the bill nobody quotes.

n8n human in the loop: Wait nodes, tool approval, and the who-clicked problem
n8n's three ways to put a person in front of an agent's action, read from its docs, and the difference between a link anyone can click and a verified approver.

Microsoft Copilot Studio approvals: what the human review actions do, and where they stop
A fair reading of Copilot Studio's Request for information action and multistage approvals, from Microsoft's docs: what they pause, who they ask, and what they leave to you.