
Shawn G
October 1, 2026
10
min. read
and updated on:
October 7, 2026
A chatbot answers. An agent acts. That is the entire distinction, and it produces two products with different costs, different risk profiles, and different failure modes.

A chatbot answers. An agent acts. That is the entire distinction, and it produces two products with different costs, different risk profiles, and different failure modes.
Most companies asking for one are describing the other. A team that says it wants a support chatbot frequently means it wants something that can process a refund, which is an agent. A team that says it wants an agent frequently means it wants better answers from its own documentation, which is a chatbot. Getting this wrong at the scoping stage is the most expensive error available in this category, because the two are not adjacent in difficulty.
| Dimension | Chatbot | Agent |
|---|---|---|
| Output | Text | Actions in other systems |
| Worst-case failure | A wrong answer | A wrong action |
| Reversibility | Inherent | Depends entirely on the action |
| Model calls per interaction | Usually one | Several |
| Reliability requirement | Good enough to be useful | High enough to be trusted |
| Typical build cost | $25,000 to $110,000 | $30,000 to $200,000 |
| Main engineering challenge | Retrieval quality | Reliability across steps |
The row that matters most is reversibility. A chatbot that gives a wrong answer is embarrassing and the user can ignore it. An agent that issues a wrong refund, sends a wrong email, or updates the wrong record has produced a consequence, and consequences require confirmation steps, audit trails, and a support process.
A chatbot performs one operation: retrieve context and generate a response. If it is right 92 percent of the time, the interaction is right 92 percent of the time, and a user who spots a bad answer moves on.
An agent performs a chain. Six steps at 95 percent each produces roughly 74 percent end-to-end success. Ten steps produces about 60 percent. Nothing is broken; each component performs well and the compounding fails.
This is why agent projects disappoint after impressive demonstrations, and it dictates the whole design approach. Fewer steps beats more. Deterministic code should handle every step that does not require judgement. And any step with a real consequence needs verification.
The user's problem is answered by information you already hold. Support questions, policy questions, documentation lookup, product guidance, and account status queries all fall here, and this covers considerably more real demand than companies expect.
The value is speed of access rather than task completion. A user who gets a correct answer in ten seconds instead of waiting two hours for a reply has been well served without anything being changed.
You want to start. A knowledge-grounded chatbot over your own content is the lower-risk entry point, and it builds the infrastructure, the server proxy, prompt versioning, cost controls, retrieval, observability, that an agent would need anyway.

The user's problem requires something to change in a system. Booking, cancelling, updating, submitting, routing, processing. Information alone does not resolve it.
The task is repetitive, high volume, and currently performed manually by someone whose time you can quantify. Agents pay back against measurable manual cost, and agents built without that comparison tend to be justified on enthusiasm.
The systems involved have real APIs. An agent operating through browser automation or a brittle interface will be dominated by the fragility of that layer, and you will spend the project's life repairing it.
And critically: the failure mode is recoverable, or a person can confirm before the action commits.
The strongest pattern for support and internal tooling is a chatbot that can hand off to a small number of narrow, confirmed actions.
It answers freely, since answering is low risk. When the user's need requires an action, it presents the action for confirmation rather than performing it: here is what I am about to do, confirm or edit. The action itself is deterministic code with a defined schema, not a model deciding freely what to call.
That design captures most of the value of an agent while keeping the risk profile close to a chatbot's, and it is why draft-then-approve is the highest-return shape in this whole category. The model prepares, a person commits.
In practice, the question that resolves this decision fastest is not what the technology can do. It is what happens if the system is wrong. If the answer is that a user reads something incorrect, build a chatbot. If the answer is that money moves or a customer receives something, build a chatbot with confirmed actions rather than an autonomous agent.
A software company wants to reduce support ticket volume. Most tickets are questions about configuration and billing policy, both documented. Verdict: chatbot, grounded in the help centre and the customer's own account state. The account-aware part is where the real value sits, because most contacts concern the customer's own situation rather than general policy. No actions needed in version one.
An insurance broker wants to process inbound claim documents. Documents arrive by email, someone reads them, extracts fields, and enters them into a system. Verdict: agent, and specifically a document processing pipeline with human approval. High volume, measurable accuracy, quantifiable manual cost, and a recoverable failure mode since a wrong extraction is caught at approval.
A marketplace wants buyers to be able to reschedule bookings conversationally. Verdict: chatbot with one confirmed action. The conversation identifies the booking and the desired time; the reschedule itself is deterministic code presented for confirmation. Building this as an autonomous agent would add risk for no gain, because the user is right there to confirm.
A logistics company wants to automate dispatch exception handling. When a delivery fails, someone decides whether to retry, reroute, or refund. Verdict: this is genuinely agentic and should still start as draft-then-approve. Let the system recommend the action with its reasoning, and a coordinator commits. Measure overturn rate for three months before considering autonomy on any single decision type.
Bolder Apps builds backends in Node.js and Laravel and sells paid discovery as a standalone engagement. The discovery deliverable that matters most on either side of this decision is the same: the step count, the consequence of an error, and the accuracy bar that constitutes done. A project specified without those three is a research programme with a delivery date attached.
A knowledge-grounded chatbot over your published content runs $25,000 to $70,000 over five to ten weeks, with retrieval quality work accounting for more of the effort than the model integration.
An account-aware chatbot, answering about the user's own data, runs $50,000 to $110,000 over eight to fourteen weeks. The premium is authentication, authorisation scoping, and the operational tooling around it.
A constrained agent handling one workflow with two or three tools and human approval runs $30,000 to $70,000 over six to ten weeks.
A multi-step production agent with several integrations, confirmation steps, observability, and traces runs $80,000 to $200,000 over twelve to twenty-four weeks.
Running cost differs more than build cost does. An agent makes several model calls per completed task rather than one, so compare cost per completed workflow rather than per request, and compare it against the manual cost it replaces rather than against a chatbot's economics.

The shared infrastructure is substantial, which is another reason to start with the chatbot and grow into agent capability.
Both need the provider call routed through your own backend, never from a client. Both need prompts as versioned artefacts rather than strings in code. Both need streaming, per-user rate limits and token caps, model tiering so simple requests do not reach your most expensive model, and caching.
Both need defined failure behaviour for timeout, rate limiting, provider outage, and unusable output. Both need moderation on input and output where users supply free text. Both need an evaluation set of representative cases run whenever prompts or models change.
And both need observability: what was sent, what came back, how many tokens, how long, per user and per feature.
Bolder Apps builds backends in Node.js and Laravel, prices project work fixed-scope rather than hourly, and is an official OpenAI partner with API credits available for qualifying projects. Worth asking any partner is which of the shared items above are in scope, because roughly 60 percent of a first AI feature is infrastructure that later features reuse, and a proposal that omits it is quoting a demo.
A chatbot needs retrieval quality. Chunking strategy, embedding choice, whether to combine vector and keyword search, and re-ranking. When a knowledge-grounded chatbot answers badly, the cause is almost always retrieval rather than the model, and teams that respond by upgrading models spend more and improve little. It also needs source display, so users can verify, and an explicit instruction and mechanism to decline when nothing relevant is retrieved.
An agent needs tool discipline and traceability. Narrowly scoped tools with precise schemas and least-privilege permissions. Validation of every parameter. Loop, token, and time limits, because a reasoning loop can loop indefinitely and the failure mode is a bill. Idempotency, so a run failing at step five can be safely resumed or abandoned. Step-level traces, because agents fail invisibly from the outcome alone. And injection resistance wherever untrusted content enters the context.
Build the chatbot first, grounded in your own content, with human handoff. Eight to twelve weeks. This establishes the shared infrastructure and, more usefully, it shows you what users actually ask for.
Then read the transcripts. The requests that could not be resolved by information are your agent requirements, defined by evidence rather than by assumption, and they are usually a much shorter list than anyone predicted.
Then add two or three of those as confirmed actions. Narrow, deterministic, with a person approving until measured accuracy justifies otherwise.
Then, only if the data supports it, remove confirmation from individual steps one at a time. Most organisations discover they never need to.
Architecturally that is close to accurate and commercially it is misleading, because the tools are where the cost, the risk, and the engineering difficulty concentrate. The conversational layer is the cheap part of an agent.
A chatbot, usually by a wide margin, because an agent makes several model calls per interaction rather than one. Model cost per completed task rather than per request when comparing, and compare an agent's cost against the manual cost it replaces rather than against a chatbot's.
Yes, and that is where support chatbots become genuinely valuable, since most contacts are about the customer's own situation. It requires authenticating the user and scoping retrieval to their data in the query itself, filtered in code rather than by instructing the model to be careful. Cross-account exposure here is a data breach, not a bug.
Frameworks help with orchestration and tracing and add abstraction that can obscure behaviour. For a first constrained agent, plain code calling the provider directly is often clearer and easier to debug. Adopt a framework when complexity justifies it.
When accuracy on that specific step has been measured against real cases over a meaningful period, the failure mode is recoverable, and monitoring exists that will tell you if performance drifts. Trust granted before measurement is how these projects generate incidents.




