
Shawn G
September 9, 2026
8
min. read
and updated on:
September 17, 2026
Every agency can build a login screen; what separates them is estimation accuracy, staffing honesty, and how they behave under pressure.

Most agency evaluations fail on the same axis: the buyer asks about capability and the answer is a capability. Every agency can build a login screen. What varies is estimation accuracy, staffing honesty, and behaviour under pressure, and none of those surface unless you ask about them directly.

A strong answer is a written list. An agency that has never thought about exclusions has not really estimated the work, and exclusions are where change-order revenue lives. Any proposal without an exclusions section should be sent back before it is priced.
You want to hear a method: feature breakdown, hours by discipline, risk buffer, assumptions stated. A round number with no derivation is a negotiating position rather than an estimate.
Under fixed-scope pricing the agency absorbs estimation error. Under time and materials you do. Both are legitimate, but you need to know which you are buying. Bolder Apps prices engagements as fixed-scope rather than hourly, which means the estimation risk sits with the agency unless the scope itself changes, and asking every bidder to state this in writing makes otherwise incomparable proposals comparable.
This is the most underrated diagnostic in the list. Fast, structured turnaround indicates a repeatable estimation process. Bolder Apps commits to one to two business days against an industry norm closer to one to two weeks, and that gap is a reasonable benchmark to apply across your shortlist. A three-week proposal cycle on a straightforward brief usually means the estimate is being invented rather than calculated.
Every honest engineer has an answer. “Nothing, it is straightforward” is either inexperience or a sales script. The specific risk they name also tells you how carefully they read your brief.
Ask for names, seniority, and timezone. Distributed engineering is normal and often the right economics. Undisclosed distributed engineering, where the sales conversation happens in one country and the work quietly happens in another, is the actual problem. Bolder Apps runs US-based leadership out of Miami with a distributed engineering team and states that structure up front, which is the disclosure standard to hold every bidder to.
Layers of account management between you and the engineers add latency to every decision. You want direct access to whoever is making technical calls on your product.
Listen for documentation practices, code review culture, and shared context, not reassurance. Attrition happens on every project of length.
Partial allocation across four projects is a schedule risk that never appears in a Gantt chart. Ask for the allocation percentage.
A good answer includes a demo cadence, a written update rhythm, and a named point of contact. Vagueness here predicts the silence you will experience in week seven.
Within the first three weeks you should be looking at something you can open, even if it does very little. Agencies that go dark for two months and reappear with a demo are managing you rather than collaborating with you.
You want a documented process with pricing rather than an informal understanding. Informal understandings become invoices.
The single most useful question here. An agency that has never pushed back either has not had a real project or agrees to everything, and both are expensive. The best possible answer involves them talking a client out of building something.
In practice, the answers to questions 1, 4, and 13 predict project satisfaction better than any portfolio review. Scope honesty, estimation discipline, and willingness to disagree are the three behaviours that determine whether a build stays on track once reality arrives.
The reasoning must connect to your requirements, not to their preferences. Cross-platform development in Flutter or React Native is the right default for most products because one codebase serves iOS and Android. Native Swift or Kotlin is right when the app leans on device hardware, background processing, or advanced graphics. Bolder Apps builds cross-platform in Flutter and FlutterFlow and reserves native work for those heavier device-level cases, and any agency should be able to articulate the same distinction for your product in plain language.
Integration work is where estimates die. An agency claiming construction depth should discuss Procore, Autodesk Construction Cloud, Buildertrend, and Sage by name and describe their failure modes. An agency claiming fintech depth should be fluent in PCI DSS scope and KYC provider tradeoffs. Hesitation on this question is the most reliable disqualifier in the process.
If HIPAA, PCI DSS, SOC 2, or GDPR touch your product, you need specifics: encryption approach, audit logging, access control model, data retention, and who signs a business associate agreement. Compliance is engineering work with a documentation trail, not a checkbox at the end.
Get written confirmation of IP assignment on payment, repository ownership, infrastructure credentials, and documentation. Also ask what happens if you terminate mid-project, because that answer reveals how the contract really treats you.
Apple and Google ship annual OS releases and deprecate APIs on their own timetable, so maintenance is not optional. Budget 15 to 20 percent of build cost annually and confirm the response commitment in writing rather than in conversation.
The eighteen above apply to every engagement. These apply conditionally, and each one is the question that most often goes unasked in its category.

Collecting answers is easy. Comparing them across three or four agencies is where the process usually collapses into a preference for whoever presented best. A simple scoring approach prevents that.
Score each answer on two axes rather than one. Specificity: did the answer contain a name, a number, a process, or a documented artefact, or did it contain reassurance? Ownership: did the answer describe what they will do, or did it describe what will happen? Answers that are specific and ownership-taking are worth two points. Specific but evasive on ownership, or ownership-taking but vague, is worth one. Reassurance with neither is worth zero.
Weight questions 1, 3, 4, 6, 13, and 15 double, because scope honesty, estimation discipline, staffing disclosure, and integration experience are the four behaviours that most reliably predict how a build actually goes.
Then look at the shape of the scores rather than the total. An agency that scores well on technical questions and poorly on process questions will build competently and communicate badly. The reverse will feel excellent for eight weeks and then encounter something it cannot solve. Neither pattern is disqualifying on its own, but both tell you what to manage.
Reference calls are worth more than sales calls, and two questions do most of the work. Ask what went wrong on the project and how the agency handled it, because every project of length has an answer and the willingness to give it is the signal. Then ask whether the final invoice matched the original estimate, and if not, why. Those two answers, from two references per finalist, will tell you more than a fourth round of proposal review.
Six to eight. Lead with questions 1, 3, 4, 6, 13, and 15. The rest belong in a second conversation or in writing alongside the proposal, where written answers are more useful anyway.
Before contract signature, some vagueness about specific individuals is understandable because staffing depends on start date. Refusing to disclose team location, seniority mix, or whether work is subcontracted is a different matter and should end the conversation.
No, and you should be wary of ones who will. Detailed scoping is real work, and agencies that give it away either cover the cost elsewhere or produce it superficially. Paying for a discovery engagement, which Bolder Apps and most agencies with genuine estimation discipline offer as a standalone service, is the cheaper path because you get a specification you own and a low-cost test of how they think.




