All Insights

    Generative AI

    Where Generative AI Actually Creates Value

    A field guide to the GenAI use cases that consistently return on investment in enterprise settings.

    Prashant BhardwajMarch 20267 min read
    Share

    Two and a half years after ChatGPT landed in the enterprise, the honest question executives are asking has changed. It used to be, "what can we do with this?" It is now, "which of the things we tried actually worked, and why do the ones we hoped would work keep disappointing us?"

    The uncomfortable answer, drawn from programmes across financial services, healthcare, retail, industrial and public-sector clients, is that a small number of use-case patterns produce most of the enterprise value. Almost everything else — including many of the demos that get boardroom applause — either underperforms or fails to reach production.

    It is worth naming the patterns that work first, and then being specific about why the popular ones that do not.

    The strongest recurring pattern is expert amplification. Take a senior specialist whose time is expensive and repetitive knowledge work is a significant fraction of their week — an underwriter, a claims examiner, a compliance analyst, a specialist doctor writing referrals, a senior engineer reviewing designs. Build a system that compresses the repetitive part of their work by half or more, and reinvest the freed capacity in higher-value activity. The math here is clean: the freed time has a known hourly cost, the reinvestment has a measurable output, and the specialist is usually the most credible advocate for further rollout.

    The second pattern is unstructured-to-structured extraction. Contracts, medical notes, incident reports, KYC documents, insurance submissions, customer complaints — enterprises are drowning in text that needs to become fields in a database. Language models are extraordinarily good at this, especially when paired with a schema, a validation step, and a human review for anything below a confidence threshold. The value is not glamorous. It is the elimination of a large, slow, error-prone data-entry cost centre.

    The third is knowledge assistance grounded in the organisation's own material. Not general chat — grounded assistance. Field engineers looking up torque values and troubleshooting steps. Contact-centre agents answering policy questions during a live call. New hires accelerating through the first six months of tacit knowledge. The pattern works when three conditions hold: the underlying knowledge is genuinely searchable and up to date; the interface is embedded in the workflow, not a separate portal; and the answers cite their source so the user can verify.

    The fourth is content and communication drafting at scale, in low-risk contexts. Marketing variants, product descriptions, internal announcements, first drafts of routine correspondence. The value here is throughput, not creativity. The pattern fails when applied to communication that carries legal or brand risk without a human review step, which is why most "AI writes our marketing" programmes end up being human-heavy in practice.

    The fifth, less discussed, is software modernisation assistance. Migrating between framework versions, translating between languages, generating tests for legacy code, explaining unfamiliar codebases to new engineers. The productivity gains here are large and measurable, and the risk profile is contained because everything the AI produces is reviewed by an engineer and executed against tests.

    Now the patterns that consistently disappoint.

    Autonomous customer-facing agents in complex domains. The demos are impressive; the production reality is that customer conversations have unusual patterns, high stakes when they go wrong, and long tails the evaluation set missed. Almost every enterprise that has tried to fully automate a customer channel has retreated to an assist-the-human model within 12 months. That model works. Full autonomy in these contexts, at the current state of the art, mostly does not.

    General-purpose enterprise chat. The pattern of "give every employee a chat window and let them figure out what to do with it" produces high initial engagement, followed by a steep drop when employees discover that the tool cannot access the systems they actually work in. The organisations getting real productivity gains from Copilot-style tools are the ones investing heavily in workflow integration and prompt libraries, not the ones treating it as a self-service novelty.

    Fully automated content pipelines without human oversight. In every industry we have seen this attempted, the failure mode is the same: the output is technically fine 95 percent of the time, and catastrophically wrong the other 5 percent, and the volume is high enough that the 5 percent produces real reputational or regulatory exposure. The economics only work with a review layer, and once the review layer is in, the productivity claims halve.

    The pattern across the wins and the disappointments is not about the technology. It is about where the AI sits in the workflow, how the residual risk is managed, and whether there is a credible unit-economics story that finance can audit. Wins are grounded, embedded, assisted, and measured. Disappointments are ungrounded, standalone, autonomous, and vague about what the baseline was.

    A practical filter for any proposed use case: can you name the human whose work is affected, the specific work unit being changed, the baseline cost or cycle time, the target after the AI is deployed, and the failure mode when the AI is wrong? If any of those five is fuzzy, the use case is not ready to build. It is ready to be scoped better.

    The organisations we see getting durable value from GenAI stopped asking "where should we deploy AI" and started asking "where do we have expensive, repetitive knowledge work with a clear baseline?" The answers are almost always specific, unglamorous, and profitable. The exciting demos are still valuable — they change what leaders think is possible — but the returns come from the boring cases done properly.

    Filed under

    GenAIROIUse Cases

    Continue reading

    More on Generative AI