The pressure to adopt is real and largely external. Vendors are attaching AI to products that did not previously have it. Staff are already using tools the organization has not evaluated. Boards are asking what the plan is. None of that pressure tells you where AI would actually help.

The organizations getting durable value share a pattern. They started from a specific operational problem and worked backward to whether AI was the right instrument. The ones that struggled started from the technology and went looking for somewhere to put it.

A tool looking for a problem produces a pilot. A problem looking for a tool produces an improvement.

The Three Categories Worth Starting With

Across most organizations, the reliable early wins fall into three groups. All three share a trait: the work is high volume, language-based, and currently consumes skilled people on unskilled tasks.

Reading and Summarizing

Work where someone reads a large volume of text to extract a small amount of signal. Reviewing submissions for completeness, pulling key terms out of contracts, summarizing correspondence into a case note, triaging incoming requests by topic.

This category works because the output is checkable. A summary can be verified against its source in less time than producing it took. That verification loop is what makes the risk manageable, and it is the reason this category tends to survive contact with a real workload.

Finding

Work where the answer already exists inside the organization but nobody can locate it. Policies, prior matters, past decisions, technical documentation, institutional knowledge sitting in documents that were never indexed for retrieval.

The value here is often larger than it first appears, because the current cost is hidden. Staff who cannot find something do not file a ticket. They ask a colleague, which consumes two people instead of one, or they recreate the work, or they proceed without it.

Prerequisite

This category depends entirely on access controls being correct first. A retrieval assistant will surface whatever it can reach, to whoever asks. If your permissions are approximate, this will find out before you do.

Drafting

Producing a structured first version that a person then edits. Standard correspondence, meeting summaries, recurring reports, initial responses to routine requests.

The economics work when editing is genuinely faster than writing, which is true for structured and repetitive output and false for anything requiring real judgment or original argument. The failure mode is subtle: a draft that is nearly right can take longer to fix than a blank page, because the reviewer has to find the errors rather than avoid them.

Where It Usually Disappoints

Three patterns account for most of the disappointing outcomes.

Common failure patterns

Evaluating a Candidate Use Case

Before committing budget or attention, work through these. If more than one answer is uncomfortable, the use case is not ready.

Question
What is the operational problem, stated without mentioning AI?

If the problem cannot be described without naming the technology, there is no problem yet. This single question eliminates most proposals that arrive from vendors.

Question
Who checks the output, and how would they know it was wrong?

A verification step that exists in principle but has no owner is not a verification step. Name the role.

Question
What data does this need to reach, and who else can reach it?

Access is the governance question that matters most and gets asked least. The tool inherits the permissions of whatever it connects to.

Question
Could you explain this use to a client, a regulator, or a board?

If the honest answer is that you would rather not, that instinct is worth listening to. It usually points at an unresolved question about data handling or accountability.

Question
What happens if it becomes unavailable next quarter?

Adoption creates dependency. Work out in advance whether the fallback is the old process, a different vendor, or nothing.

A Note on Sequencing

The instinct is to run a broad pilot and see what sticks. Broad pilots tend to produce broad ambivalence. A single use case, owned by a named person, measured against the process it replaced, tells you more in six weeks than a general rollout tells you in a year, and it builds the internal judgment you need before the second one.

Where to start

Identify one high-volume, language-heavy task where a person currently reads a lot to extract a little, and where a knowledgeable reviewer already sees the output. That combination gives you value and a safety net in the same workflow.