
Shawn G
September 29, 2026
9
min. read
and updated on:
October 5, 2026
Generative features produce new content rather than classifying or extracting existing content, and that distinction changes the engineering problem, the cost structure, and the legal exposure. Text...

Generative features produce new content rather than classifying or extracting existing content, and that distinction changes the engineering problem substantially. Classification has a right answer you can test against. Generation has a distribution of acceptable outputs, no objective correctness, and a set of legal and reputational exposures that classification does not carry.
Those exposures are the reason generative features fail in production more often than their demos suggest, and none of the causes are about model quality.
Text. The most mature and the cheapest. Drafting, rewriting, summarising, translating, expanding. Latency is manageable with streaming, cost is predictable per token, and quality is assessable by a human in seconds. This is where almost every product should start.
Images. Cost per generation is higher and fixed rather than proportional to length, latency runs seconds rather than milliseconds, and quality control is genuinely hard because subtle defects are common and users notice them. Requires a review or regeneration flow rather than direct publication.
Voice and speech. Synthesis is fast and inexpensive enough for real use, and voice cloning carries consent and likeness obligations that text does not. Transcription in the other direction is mature and reliable.
Video. Expensive, slow, and still inconsistent. Viable for short clips in specific use cases and rarely justifiable as a core feature in a commercial product on current economics.
The pattern that works is replacing something a user currently writes or produces from scratch, with the user reviewing before use.
Concretely: product descriptions from specifications and images, which is a genuine time saving for large catalogues. Draft replies in support and sales tooling. First-draft documentation and reports from structured data. Marketing variants for testing. Listing content in marketplaces, where completion rates measurably improve. Meeting summaries from transcripts.
Each of these has a common shape worth noticing: the output is a draft, a human edits it, and the value is the blank page being filled rather than the work being finished. Features designed as drafts succeed. Features designed to replace judgement generate incidents.
Any feature that generates content from user input needs moderation on both sides, and any feature that publishes generated content needs it more.
Screen input for prompts attempting to produce harmful or infringing output. Screen output before display or publication. Provider moderation endpoints handle a large share of this, and category thresholds need tuning to your context rather than accepting defaults.
Both app stores require content reporting mechanisms and moderation for apps generating or hosting content, and apps that let users produce and share generated content receive closer review scrutiny. This is a submission requirement rather than a best practice, and apps that omit it get rejected.
Where generated content reaches other users, add a reporting route, a review queue, and a takedown path. This is the operational tooling that makes a generative feature safe to run, and it is invisible in every brief.
Four issues, and none of them are engineering problems.
Who owns the output. Provider terms vary and generally assign output rights to the user or customer. Your own terms need to say what you claim and what your users hold, and it needs to be consistent with your provider agreement.
Whether your data trains the model. Enterprise and business API tiers typically commit to not training on submitted data and to defined retention. Consumer tiers frequently do not. This is a question your customers will ask, particularly business customers, and the answer belongs in your documentation.
Likeness and voice. Generating a recognisable person’s likeness or voice raises publicity and consent issues that vary by jurisdiction and are moving quickly. Voice cloning in particular requires documented consent from the speaker.

Disclosure. Requirements to disclose AI-generated content are expanding across jurisdictions, and disclosure is also the trust-preserving choice regardless. Users who know a response is generated read it appropriately, and concealment tends to be discovered.
In practice, the constraint that most often stops a generative feature from shipping is not quality. It is that nobody worked out who reviews the output before it reaches a customer. Build the review step first and the quality problem becomes manageable.
Generative features fail in the interface as often as in the model, because they break an assumption users hold about software: that the same input produces the same output.
Four design decisions carry most of the outcome.
Set expectations before generating. Label the output as a draft. Users who understand they are receiving a starting point edit it. Users who expect a finished answer judge it as one and find it wanting.
Make editing the primary action. The generated content should land in an editable field rather than in a display panel with a copy button. The interaction you want is refinement, not accept or reject.
Make regeneration cheap and obvious. Users will not take the first output. A visible regenerate control with some variation is the mechanism that makes the feature feel useful rather than unreliable.
Show progress honestly for slow modalities. Image and video generation take long enough that a spinner is not enough. Queue position, estimated time, or a notification on completion. Users tolerate waiting when they know they are waiting.
The pattern to avoid is a single button that produces one output with no path to improve it. That design makes every imperfect generation a dead end, which is how features with acceptable quality get abandoned as bad.
Six questions to answer before development, because each one changes the estimate materially.
Which modality, and does the latency profile suit a synchronous request or require queuing and notification?
Who reviews the output before it reaches a customer or another user, and what tooling do they need?
What is the quota per user per period, and does it differ by pricing tier? This is a product decision with direct cost consequences and it should not be deferred.
What happens to generated assets, meaning storage, retention, user deletion, and whether you retain copies for quality review?
What does moderation cover, on input and on output, and where does flagged content go?
What is the accuracy or acceptance bar that constitutes done? Without a number, quality improvement is unbounded and the project has no end.

Bolder Apps is an official OpenAI partner, prices project work fixed-scope rather than hourly, and sells paid discovery as a standalone engagement. On generative features specifically the discovery deliverable that earns its cost is the answer to the last two questions, because a generative feature scoped without a review path and an acceptance bar becomes a research programme with an invoice attached.
| Modality | Cost character | Practical control |
|---|---|---|
| Text | Per token, proportional to length | Model tiering, context limits, caching |
| Images | Per generation, fixed and higher | Generation quotas, resolution tiers, caching |
| Voice synthesis | Per character or second | Caching repeated phrases |
| Video | Per second, high | Hard quotas, paid tier only |
Image and video generation change the economics of a free tier fundamentally. Text features can often absorb a generous allowance; image generation at a few cents per attempt, with users regenerating four times to get one they like, cannot. Model cost per active user before building and design the quota into the product rather than adding it after a bill arrives.
Bolder Apps is an official OpenAI partner with API credits available for qualifying projects and prices project work fixed-scope rather than hourly. On generative work the fixed scope matters because output quality improvement is unbounded, and an agreed quality bar is what turns the feature into a delivery rather than a standing workstream.
Beyond the standard architecture of a server-side proxy, versioned prompts, streaming, and cost controls, generative features need four additional things.
Asynchronous handling for slow modalities. Image and video generation take long enough that a synchronous request is the wrong shape. Queue the job, return an identifier, notify on completion.
Regeneration as a first-class flow. Users will not accept the first output. Cheap regeneration with variation, and a way to keep or discard, is the interaction pattern that makes generative features usable.
Storage and retention decisions. Generated assets accumulate quickly and cost money to store. Decide retention policy, whether users can delete, and whether you retain for quality review, before launch.
Provenance metadata. Recording that an asset was generated, by which model and prompt version, matters for disputes, for disclosure obligations, and for debugging quality regressions.
Generation has no correct answer, which does not mean it cannot be measured.
Track acceptance rate, meaning how often users keep the first output. Track edit distance, meaning how much they change it, which is the strongest available proxy for usefulness on text. Track regeneration count, since three attempts per usable result is a cost and quality signal simultaneously. And sample outputs for human review on a schedule.
Then build a small evaluation set of representative inputs and run it whenever you change a prompt or model, so quality changes are observed rather than assumed. Without this, prompt changes are guesses validated by whichever outputs someone happened to read.
Usually not first. Prompt design, better context, and structured constraints solve most quality problems far more cheaply. Fine-tuning earns its cost for consistent formatting, a specific brand voice, or narrow domain output where prompting has demonstrably plateaued.
Technically yes, and you should think hard before allowing it. Direct publication without review means your product’s name is attached to whatever the model produces, including its failures. A review step, even a lightweight one, converts a reputational risk into a productivity feature.
Treat it as a real risk rather than a theoretical one, particularly for images and for style imitation of identifiable creators. Moderation on prompts, terms of use that place responsibility appropriately, and a takedown process are the practical mitigations, and legal advice is warranted if this is central to your product.
Routinely, provided moderation exists, content reporting is available, and privacy disclosures accurately reflect data sent to third parties. Apps generating content in sensitive categories receive closer scrutiny, and apps with no moderation get rejected.
Per-user quotas enforced server-side, input and output moderation, rate limiting, and monitoring for unusual per-account consumption. All four require the provider call to route through your own backend, which it should regardless.




