Skip to content

Practitioner Guide: Running a Workflow Optimization Program

What this guide is for

This guide is written for the operations leader, transformation lead, or workflow program owner who has been handed a mandate that sounds like "use AI to make our processes faster." That mandate arrives in many forms — a board directive to "do something with AI," a cost-reduction target, a backlog of teams asking for automation, or the uncomfortable realization that competitors are moving faster. All of those converge on the same underlying work: deciding which workflows to change, redesigning them so AI actually helps rather than merely sits on top, moving the people who run those workflows through the transition, and proving the change paid off.

This guide assumes you are not starting from zero on AI capability. The strategy, governance, platform, and data foundations are covered in other tracks; this guide is about the operational discipline of running an optimization program against real workflows. It is deliberately opinionated, because the failure modes here are well documented and expensive. The dominant lesson from a decade of automation programs is not that the technology fails — it is that organizations automate the wrong things, automate them badly, fail to bring people along, and never measure whether anything improved.

What this guide produces is a runnable program in four phases: a prioritized portfolio of workflow opportunities backed by evidence rather than opinion (Phase 1), redesigned workflows with humans deliberately kept in the loop (Phase 2), a change effort that gets the people doing the work to actually adopt the new way (Phase 3), and an instrumented measurement loop that tells you whether to scale, fix, or kill each initiative (Phase 4). The phases are sequential per workflow but run as a rolling pipeline across a portfolio — you are always discovering the next batch while building and measuring the current one.

The stakes are high enough to state plainly up front. A 2025 MIT study of enterprise generative-AI deployments found that roughly 95% of pilots delivered little to no measurable impact on profit and loss, with only about 5% capturing substantial value (MIT NANDA, 2025). The gap between those two groups is almost never the model. It is the program discipline this guide is about.


Phase 1: Discovery and prioritization (weeks 1–4)

The goal of Phase 1 is a prioritized portfolio of workflow opportunities, each with a documented baseline and a clear hypothesis about what AI changes and why. The single most common reason optimization programs fail is that they skip this phase — they start with a tool or a flashy use case and reverse-engineer a justification. You want the opposite: evidence first, solution second.

Why discovery is worth four weeks

Most organizations badly underestimate how much of their actual work is invisible. Processes drift from their documented versions, exception paths multiply, and a surprising share of effort goes to work nobody designed. McKinsey's foundational automation research found that about half of all work activities are technically automatable with demonstrated technology, and that roughly 60% of occupations have at least 30% of their constituent activities automatable — meaning the opportunity is distributed across tasks within roles, not concentrated in whole jobs you can simply eliminate (McKinsey, 2017). Generative AI raised that ceiling considerably: McKinsey later estimated that current technologies including gen AI could automate activities absorbing 60–70% of employees' time today (McKinsey, 2023). The point of discovery is to find where, in your specific organization, that potential actually lives.

There is real waste to recover. Knowledge workers spend roughly a fifth of the workweek — about 19% — just searching for internal information (McKinsey, 2012). That is the kind of diffuse, hard-to-see drain that disciplined discovery surfaces and a redesigned workflow can attack.

Run a discovery sprint

Structure discovery as a focused sprint, not an open-ended assessment that drags for a quarter. The deliverable is a ranked opportunity list, and the sprint should produce it in three to four weeks.

  • Frame the scope. Pick one or two business areas (a department, a function, an end-to-end process such as order-to-cash). A program that tries to discover everything at once discovers nothing usefully.
  • Go to the work. Interview five to ten people who actually run the workflows, and watch them work. Ask what slows them down, where they wait, what they redo, and which tasks they would happily never do again. The friction people complain about is your richest signal.
  • Map the current state. For each candidate workflow, document the steps, the handoffs, the systems touched, the volume (how many times per day/week), the cycle time, the error/rework rate, and the people involved. This map is also your baseline — capture the numbers now, because once you change the workflow you will never recover a clean "before."

Use process and task mining where the volume justifies it

For high-volume, system-mediated workflows, do not rely on interviews alone — interviews capture what people think they do, not what the event logs show actually happens. Process mining (reconstructing the real process from system event logs) and task mining (capturing desktop-level actions) reveal the hidden variants, rework loops, and bottlenecks that nobody describes in an interview. The discipline has gone mainstream, and it measurably accelerates discovery — in Deloitte's intelligent-automation research, 63% of respondents said process intelligence sped up the discovery of automation use cases, versus 7% who said it slowed them down (Deloitte, 2022).

Discovery is not optional overhead. In Deloitte's research, 63% of organizations reported that process intelligence accelerated the discovery of automation opportunities — and the firms that skip discovery are precisely the ones that end up automating processes they never understood (Deloitte, 2022).

Prioritize with a value/feasibility lens

Once you have candidates with baselines, score them. A simple two-axis model is enough to start: value (volume × time saved per instance × cost per hour, plus quality and risk-reduction benefits) against feasibility (data availability, process standardization, technical complexity, and change difficulty). Plot the candidates, and pursue the high-value/high-feasibility quadrant first.

Two prioritization rules earn their place from hard experience. First, deprioritize any workflow that is not yet standardized. Automating a chaotic process locks in the chaos — more on this in the failure-patterns section. Second, weight change difficulty explicitly. A technically easy automation in a politically resistant team will fail; a moderately hard one in a willing team will succeed. Feasibility is as much organizational as technical.

Phase 1 deliverable: a prioritized opportunity portfolio. For each selected workflow: a current-state map, a documented quantitative baseline (volume, cycle time, error/rework rate, cost), a value/feasibility score, and a one-paragraph hypothesis stating what AI will change and the expected outcome. Pair this with the Workflow Maturity & Opportunity Scoring assessment to pressure-test your ranking.


Phase 2: Redesign and build with AI in the loop (weeks 5–10+)

The goal of Phase 2 is a redesigned workflow — not the old workflow with an AI step bolted on. This distinction is the single biggest determinant of whether the program delivers value, and the data on it is unusually clear.

Don't pave the cow path

The instinct under deadline pressure is to take the existing process and insert AI wherever a step looks automatable. This is "paving the cow path" — formalizing an inefficient route instead of asking whether the route should exist at all. The warning is as old as business computing: as Bill Gates put it, "automation applied to an efficient operation will magnify the efficiency [and] automation applied to an inefficient operation will magnify the inefficiency" (Gates, The Road Ahead, 1996).

The modern evidence is striking. In McKinsey's 2025 State of AI survey, among 25 organizational attributes tested, fundamentally redesigning workflows had the single biggest effect on whether an organization saw EBIT impact from generative AI, and AI high performers were far more likely to have redesigned workflows rather than layering AI onto existing ones (McKinsey, 2025). Yet only about 21% of organizations using gen AI said they had fundamentally redesigned even some workflows (McKinsey, 2025) — meaning roughly four in five were bolting AI on without rethinking the work. The redesign gap is exactly where the value gap comes from.

So the first redesign question is never "where do we insert AI?" It is "if we were designing this workflow today, knowing what AI can now do, how would the work flow?" Then build toward that target state. Eliminate steps that exist only to compensate for a constraint AI removes. Collapse handoffs. Move quality checks earlier.

Keep the human in the loop — by design, not by accident

Full autonomy is rarely the right starting point, and the research on AI augmentation explains why. Where AI augments skilled humans, the gains are large and real:

Study Setting Result
Brynjolfsson, Li & Raymond — Generative AI at Work (NBER, 2023) Customer-support agents +14% issues resolved per hour on average; +34% for novice agents (source)
Dell'Acqua et al. — Navigating the Jagged Technological Frontier (HBS/BCG, 2023) Management consultants, tasks within AI's frontier +12.2% tasks completed, 25.1% faster, >40% higher quality (source)
GitHub — Quantifying Copilot's impact (2022) Software developers 55% faster on a defined coding task (source)

But the same consulting study contains the warning: on a task deliberately chosen to fall outside AI's competence, consultants using AI were 19 percentage points less likely to reach the correct answer than those working without it (Dell'Acqua et al., 2023). AI applied confidently to the wrong task degrades performance. That is the entire case for designing the human checkpoint deliberately rather than hoping people will catch errors.

Practical human-in-the-loop design:

  • Match the integration level to the stakes. Low-risk, reversible tasks can run with light review; high-stakes or irreversible decisions need an explicit human approval gate. The Four Levels of Workflow AI Integration gives you a vocabulary for choosing the right level per step.
  • Make review meaningful, not rubber-stamp. Surface the AI's confidence and its reasoning so the human is reviewing a judgment, not just clicking approve. Route only the uncertain cases to humans where you can; reserve scarce expert attention for where it changes the outcome.
  • Design the fallback. Decide explicitly what happens when the AI is unavailable, slow, or low-confidence — does the work stop, proceed with a default, or escalate? Untested fallbacks are a top source of production incidents.

Build, pilot, and instrument from day one

Build the redesigned workflow as a contained pilot with real users and real volume — not a demo on cherry-picked inputs. Instrument it before launch: the metrics you will judge it on (cycle time, throughput, quality, cost, adoption) must be wired in from the first day, against the Phase 1 baseline. A pilot you cannot measure is a pilot you cannot defend.

Refer to the Tooling Landscape for build options, and confirm the underlying data meets the bar in Data Readiness — redesigns routinely stall because the data the new workflow assumes isn't actually available or clean.

Phase 2 deliverable: a redesigned, instrumented workflow running in a controlled pilot with real users; documented human-in-the-loop checkpoints and fallback behavior; and a target-state process map showing what changed from the Phase 1 current state and why.


Phase 3: Change management and rollout (overlapping, weeks 8–16+)

The goal of Phase 3 is adoption — the redesigned workflow becoming how people actually work, not a tool that sits unused while everyone quietly reverts to the old way. This phase is where most of the program's risk lives, and it is the phase most often under-resourced.

Treat adoption as the hard part, because it is

The base rate for organizational change is sobering. Across fifteen years of McKinsey research on transformations, fewer than one-third succeed at both improving performance and sustaining the gains (McKinsey, 2021) — the empirical backbone behind the familiar "most transformations fail" line (the round 70% figure traces to Kotter). AI does not get a pass on this; it inherits the same adoption physics, plus a layer of fear specific to automation.

That fear is measurable. A 2025 Pew survey found that 52% of U.S. workers were worried about the future impact of AI in the workplace, against just 36% who felt hopeful, and 32% expected it to reduce their job opportunities (Pew Research Center, 2025). You cannot manage that away with a launch email. People who fear a tool will replace them do not adopt it well.

The flip side is that change management is the most reliable lever you have. Prosci's research across more than a decade of data finds that initiatives with excellent change management are seven times more likely to meet objectives than those with poor change management — roughly 88% meeting or exceeding objectives versus 13% (Prosci, 2023). The investment pays for itself.

Run change management in parallel, not at the end

Start the change effort during Phase 2 build, not after launch. The core moves:

  • Involve the people who do the work in the redesign. The frontline staff you interviewed in discovery should help shape the new workflow. Co-design converts the most credible potential resisters into advocates and surfaces practical failure modes engineers miss.
  • Be honest and specific about the role change. Vague reassurance ("AI will augment, not replace you") breeds distrust. Say concretely what the AI takes over, what the human now owns, and how the role changes. The augmentation framing is only credible if the redesign actually reflects it.
  • Invest in reskilling early. The skills demand is real and large: the World Economic Forum projects that 59% of the global workforce will need training by 2030 (WEF, 2025), and that 39% of workers' core skills will be transformed or outdated over 2025–2030 (WEF, 2025). Pair every significant workflow change with the specific capability-building it requires; coordinate with Talent & Capability Building.
  • Close the leadership-perception gap. Leaders systematically underestimate how much their people are already using AI — McKinsey found 13% of employees use gen AI for at least 30% of their daily work, while the C-suite estimated just 4%, a roughly 3x gap (McKinsey, 2025). Your workforce is likely further along (and using more shadow tools) than leadership assumes. Build on that existing momentum instead of treating adoption as a standing start.

Roll out in waves with reinforcement

Expand from pilot to full rollout in waves, not a single cutover. Each wave: train the cohort, run the new workflow alongside support, gather feedback, fix the friction, then expand. Reinforcement matters — adoption decays without it. Make the new way easier than the old way (remove access to the legacy path once a team is live), celebrate the early adopters, and keep a visible feedback channel so problems surface before they become reasons to revert. Broader adoption and culture practices live in AI Adoption & Culture.

Phase 3 deliverable: the redesigned workflow adopted by the target population, with documented training, a defined role-change communication, an active feedback loop, and adoption metrics (active users, share of volume running through the new workflow) trending toward target.


Phase 4: Measurement and iteration (continuous, from launch)

The goal of Phase 4 is a decision: scale this workflow, keep fixing it, or kill it — made against instrumented evidence rather than anecdote. Measurement is not a closing report; it is a continuous loop that starts the day the pilot goes live.

Measure, because most organizations don't

The measurement gap is the quiet killer of optimization programs. McKinsey's 2025 survey found that more than 80% of organizations reported no tangible enterprise-level EBIT impact from gen AI (McKinsey, 2025), and fewer than 20% tracked well-defined KPIs for their gen AI initiatives (McKinsey, 2025). The two facts are connected: you cannot manage toward an impact you never instrument. Independent surveys echo it — 46% of organizations report having no structured ROI measurement framework at all (Wavestone, 2025), and Deloitte found 41% of organizations struggled to define and measure the impact of their gen AI efforts (Deloitte, 2024), with only 16% producing regular value reports for the CFO (Deloitte, 2024).

Fewer than 20% of organizations track well-defined KPIs for their generative-AI initiatives (McKinsey, 2025), and more than 80% report no enterprise-level EBIT impact (McKinsey, 2025). Those two numbers are not a coincidence — unmeasured value is, in practice, uncaptured value.

When workflows are redesigned and measured, the results justify the effort. McKinsey documented a North American telecom that paired gen AI with workflow redesign in customer care and cut total call volume by about 30%, reduced average handle time by more than 25%, and lifted first-call resolution by 10–20 percentage points (McKinsey, 2024). And among organizations that do measure, the ROI lands: nearly three-quarters report their most advanced gen AI initiative is meeting or exceeding ROI expectations (Deloitte, 2024).

What to instrument

Measure against the Phase 1 baseline. Four dimensions, plus adoption:

Dimension What to instrument Watch for
Cycle time / throughput End-to-end time per instance; volume processed per period Faster steps that don't shorten the end-to-end time (the bottleneck moved, not disappeared)
Quality Error rate, rework rate, exception rate, downstream complaint/defect rate Speed gains paid for with quality loss; AI errors slipping past review
Cost Fully loaded cost per instance (labor + AI/compute + tooling) AI/compute and oversight cost erasing the labor savings
Adoption Active users; share of total volume running through the new workflow High deployment, low usage — the workflow exists but people revert
Outcome / value The business metric the workflow ultimately serves (revenue, retention, SLA) Local optimization that doesn't move the metric that matters

Two instrumentation disciplines separate credible programs from theatre. Baseline before you change — Phase 1 captured this; without it your post-deployment numbers are unanchored. And measure net, not gross — the savings that count are after AI/compute costs and the cost of human oversight. A workflow that "saves" twenty hours of labor but adds twenty-five hours of review and a large inference bill is a loss wearing a win's clothing.

Run the iteration loop and make the kill decision

Review each workflow on a fixed cadence (monthly is reasonable early on). Compare to baseline and target. Then act: where the workflow beats target, scale it to the next population; where it underperforms but the cause is fixable (a redesign flaw, a data gap, a training gap), fix and re-measure; where it has been given a fair run and the value isn't there, kill it. The willingness to kill underperformers is a feature, not a failure — it is precisely what the ~5% of organizations capturing real value do that the rest don't. Feed the consolidated results into enterprise value tracking via Measurement & Value Realization.

Phase 4 deliverable: a live measurement dashboard per workflow (baseline vs. current across the dimensions above), a fixed review cadence, and a documented scale/fix/kill decision for each initiative.


Common failure patterns and how to avoid them

These are the patterns that recur across automation and AI programs. Each one is avoidable; each one is common precisely because the path of least resistance leads straight into it.

Automating a bad process. The most expensive mistake. Lift-and-shift automation of an unstandardized, inefficient workflow magnifies the inefficiency (Gates, 1996) and locks it in. Avoid it: standardize and redesign before you automate; deprioritize any candidate that isn't yet stable (Phase 1).

Bolting AI on instead of redesigning. The data is unambiguous: workflow redesign had the single largest effect on EBIT impact in McKinsey's 2025 study, yet only ~21% of gen-AI users had done it (McKinsey, 2025). Inserting an AI step into the old flow captures a fraction of the available value. Avoid it: design the target-state workflow first (Phase 2).

The pilot that never scales. The signature failure of the era. Roughly 95% of enterprise gen-AI pilots delivered no measurable P&L impact (MIT NANDA, 2025); Gartner predicted at least 30% of gen-AI projects would be abandoned after proof of concept by end of 2025 (Gartner, 2024); and the share of companies scrapping the majority of their AI initiatives jumped from 17% to 42% year over year (S&P Global Market Intelligence, 2025). Pilots stall because they were never designed to scale, never measured, or never adopted. Avoid it: instrument from day one (Phase 2), run real change management (Phase 3), and gate scaling on measured results (Phase 4).

The scale gap, learned the RPA way. This is not new. EY observed that 30–50% of initial RPA projects failed (EY, 2016), and Deloitte found only 3% of organizations had scaled RPA to 50+ bots (Deloitte, 2018). The cause then is the cause now: programs that don't standardize processes and don't build the operating discipline to run automations at scale. Avoid it: treat the program as a rolling portfolio with governance, not a series of one-off projects.

Ignoring the people. Fewer than one-third of transformations succeed (McKinsey, 2021), and over half of workers are worried about AI at work (Pew, 2025). A technically perfect workflow that people fear or resent will not be adopted. Avoid it: co-design with frontline staff, communicate role changes honestly, reskill early, and reinforce after launch (Phase 3).

Flying blind on value. Fewer than 20% of organizations track well-defined KPIs (McKinsey, 2025) and 46% have no ROI framework (Wavestone, 2025). Without a baseline and net measurement, you cannot tell the wins from the losses, so you scale both. Avoid it: baseline in Phase 1, instrument in Phase 2, and measure net value continuously in Phase 4.


What good looks like

A well-run workflow optimization program, a few quarters in, looks like this:

Discovery is evidence-led. Workflow selection is driven by documented baselines and value/feasibility scoring — not by whichever use case had the loudest sponsor. For high-volume processes, the selection reflects what the event logs show, not just what people said in interviews.

Workflows are redesigned, not paved. Each shipped workflow looks materially different from its predecessor — steps eliminated, handoffs collapsed, checks moved earlier — because the team asked how the work should flow given AI, not where to insert a model.

Humans are in the loop on purpose. Review checkpoints sit where the stakes justify them, route the uncertain cases to people, and surface the AI's reasoning so review is real. Fallback behavior is defined and tested.

People adopted the change. The share of volume running through the new workflow is high and rising, the legacy path has been retired for live teams, and the frontline talks about the new workflow as theirs because they helped design it. Reskilling kept pace with the role changes.

Every workflow is measured net, against a baseline. A dashboard shows cycle time, quality, cost, adoption, and the business outcome — before versus now — with AI/compute and oversight costs netted out. The monthly review produces a clear scale/fix/kill decision, and underperformers actually get killed.

The program runs as a portfolio. Discovery, build, rollout, and measurement run as a rolling pipeline, with governance and a standard operating rhythm — so the organization is in the ~5% that captures real value rather than the majority whose pilots quietly evaporate.


Named frameworks you'll hear pitched for this

Consultancies and platform vendors have started publishing their own named methodologies for the exact program this guide describes, each with a memorable acronym. Read against the four phases above, they are a relabeling, not an addition. Audit/Assess and Gauge map onto Phase 1's discovery and the opportunity-scoring assessment. Engineer maps onto Phase 2's redesign and build. Navigate maps onto the human-oversight axis in The Four Levels of Workflow AI Integration. Track maps onto Phase 4's measurement loop. Neither acronym has a step for Phase 3 — change management and rollout — which this guide argues is where most of a program's real risk lives.

Two of these are worth knowing by name, and worth not confusing with each other: DAIN Studios and ServiceNow have each independently published a framework called A.G.E.N.T. — same acronym, different letters, different companies (DAIN Studios, 2026; ServiceNow, 2026). That collision is itself a tell: when a mnemonic conveniently spells the word being sold, the letters were chosen to fit the acronym, not the other way around. ServiceNow's version adds one genuinely concrete detail DAIN's doesn't — a target accuracy rate before advancing a pilot past the "75-80%" mark — though treat that as directional: a single accuracy scalar is a loose fit for a multi-step agent, where what actually needs measuring is per-step tool selection and full-trajectory completion, not one number with an unstated denominator.

If you encounter DAIN's version specifically: it's the methodology behind a paid executive course, Harvard Data Science Initiative's "Agentic AI Intensive," run as "a collaboration between Next Gen Learning and Harvard Data Science Initiative" and drawing on scholarship from the separate Harvard Data Science Review journal (HDSI/Next Gen Learning, 2026). Two of the course's six named instructors are DAIN's own co-founders. That makes it a real collaboration with genuinely independent faculty alongside them — not Harvard independently vetting and endorsing DAIN's framework.

What neither acronym closes is the set of problems that actually break agent pilots in week one. Evaluation — a multi-step agent can't be scored the way a single classifier can; you need trajectory-level scoring (did it call the right tools, in a defensible order, recoverably when a step failed), not just a final-answer accuracy number. Reliability versus safety, which are different problems: non-determinism (the same input yields different runs, requiring consistency metrics and eval tolerance) is not the same problem as irreversible actions (sending an email, issuing a refund), which needs confirmation gates, dry-run modes, and rollback logic regardless of how deterministic the model is. Error compounding — a step that's 95% reliable on its own is roughly 60% reliable chained ten steps deep , which is more often why an agent demo falls apart at scale than any single step being unreliable. Context and tool-selection limits — less about raw context-window size now than about retrieval quality degrading over long trajectories and tool-selection accuracy dropping as the available toolset grows, which is the actual argument for decomposing one large agent into scoped sub-agents rather than building one that does everything. All four are architecture and evaluation problems, covered in Technology Architecture & Platform, not in either acronym.


For the conceptual model behind this program, see the Workflow Optimization Framework. To choose the right degree of automation per step, see The Four Levels of Workflow AI Integration. For build options, see the Tooling Landscape. To score and rank opportunities, work through the Workflow Maturity & Opportunity Scoring assessment.

Sources

  • MIT Project NANDA — The GenAI Divide: State of AI in Business 2025, 2025. Just 5% of integrated AI pilots are extracting millions in value, while the vast majority remain stuck with no measurable P&L impact. View source · verified 2026-06-20 · ⚠ secondary mirror
  • McKinsey Global Institute — A Future That Works: Automation, Employment, and Productivity, 2017. about half the activities people are paid to do globally could theoretically be automated... fewer than 5 percent of occupations can be entirely automated, [but] about 60 percent have at least 30 percent of activities that could be automated. View source · verified 2026-06-21 · ⚠ secondary mirror
  • McKinsey (McKinsey Global Institute) — The economic potential of generative AI: The next productivity frontier, 2023. the acceleration in the potential for technical automation is largely due to generative AI's increased ability to understand natural language, which is required for work activities that account for 25 percent of total work time. View source · verified 2026-06-20 · ⚠ secondary mirror
  • McKinsey Global Institute — The Social Economy: Unlocking Value and Productivity Through Social Technologies, 2012. interaction workers... spend an estimated 19 percent of their time searching and gathering information. View source · verified 2026-06-21 · ⚠ secondary mirror
  • Deloitte — Automation with intelligence (Global Intelligent Automation survey), 2022. 63 per cent of respondents believe that process intelligence accelerated the discovery process... just seven per cent believe that it slowed down or stopped discovery. View source · verified 2026-06-22 · primary
  • McKinsey — The State of AI, 2025. 21% of respondents reporting gen AI use say their organizations have fundamentally redesigned at least one workflow; redesigning workflows has the biggest effect on an organization's ability to see EBIT impact from gen AI. View source · verified 2026-06-20 · primary
  • Brynjolfsson, Li & Raymond — Generative AI at Work (NBER Working Paper 31161), 2023. Access to the tool increases productivity, measured by issues resolved per hour, by 14% on average, including a 34% improvement for novice and low-skilled workers but with minimal impact on experienced and highly skilled workers. View source · verified 2026-06-20 · primary
  • Dell'Acqua et al. — Navigating the Jagged Technological Frontier (HBS Working Paper 24-013), 2023. In a pre-registered experiment with 758 consultants, those using AI inside the frontier completed 12.2% more tasks, 25.1% more quickly, with more than 40% higher quality; on a task outside the frontier AI users were 19 percentage points less likely to be correct. View source · verified 2026-06-20 · primary
  • Peng et al. — The Impact of AI on Developer Productivity: Evidence from GitHub Copilot, 2023. The treatment group with access to the AI pair programmer completed the task 55.8% faster than the control group; developers with GitHub Copilot took on average 1 hour 11 minutes versus 2 hours 41 minutes, across 95 professional programmers. View source · verified 2026-06-20 · primary
  • McKinsey — The science behind successful organizational transformations, 2021. Less than one-third of respondents... say their companies' transformations have been successful at both improving organizational performance and sustaining those improvements over time. View source · verified 2026-06-21 · ⚠ secondary mirror
  • Pew Research Center — U.S. Workers Are More Worried Than Hopeful About Future AI Use in the Workplace, 2025. About half of workers (52%) say they are worried about the future impact of AI use in the workplace; 36% say they feel hopeful; about a third (32%) say it will lead to fewer opportunities for them. View source · verified 2026-06-20 · primary
  • Prosci — The Correlation Between Change Management and Project Success, 2023. Percent meeting or exceeding project objectives: excellent change management 88%, good 73%, fair 39%, poor 13% (roughly seven times more likely with excellent vs poor change management). View source · verified 2026-06-20 · primary
  • World Economic Forum — The Future of Jobs Report 2025, 2025. 59% of the global workforce is projected to require reskilling or upskilling by 2030... 11 of whom are unlikely to receive it. View source · verified 2026-06-21 · primary
  • World Economic Forum — Future of Jobs Report 2025, 2025. On average, workers can expect that two-fifths (39%) of their existing skill sets will be transformed or become outdated over the 2025-2030 period; down from 44% in the 2023 edition. View source · verified 2026-06-20 · primary
  • McKinsey — Superagency in the Workplace, 2025. C-suite leaders estimate that only 4 percent of employees use gen AI for at least 30 percent of their daily work, [but] 13 percent of employees report that they use gen AI at that level. View source · verified 2026-06-21 · ⚠ secondary mirror
  • McKinsey — The State of AI, 2025. More than 80 percent of respondents say their organizations aren't seeing a tangible impact on enterprise-level EBIT from their use of gen AI. View source · verified 2026-06-21 · ⚠ secondary mirror
  • McKinsey — The State of AI: How Organizations Are Rewiring to Capture Value, 2025. Less than one in five organizations are tracking KPIs for gen AI solutions — and tracking well-defined KPIs is the practice with the most impact on the bottom line. View source · verified 2026-06-20 · primary
  • Wavestone — Global AI Survey 2025: The paradox of AI adoption, 2025. while most organizations see tangible benefits, 46% do not yet have a structured ROI measurement framework. View source · verified 2026-06-21 · primary
  • Deloitte — The State of Generative AI in the Enterprise: Now decides Next (Q3), 2024. 41% have struggled to define and measure the exact impacts of their GenAI efforts. View source · verified 2026-06-20 · primary
  • Deloitte — The State of Generative AI in the Enterprise: Now decides Next (Q3), 2024. Only 16% have produced regular reports for the CFO about the value being created with GenAI. View source · verified 2026-06-20 · primary
  • McKinsey — From promising to productive: Real results from gen AI in services, 2024. Total call volume fell by about 30 percent, and average handle time by more than one-quarter, even as service quality improved: first-call resolution rates rose by ten to 20 percentage points. View source · verified 2026-06-21 · ⚠ secondary mirror
  • Deloitte — State of Generative AI in the Enterprise (Q4 / Wave 4: Generating a New Future), 2024. nearly three-quarters of respondents reported that their most advanced GenAI initiative is meeting or exceeding ROI expectations. View source · verified 2026-06-22 · primary
  • Gartner — Gartner Predicts 30% of Generative AI Projects Will Be Abandoned After Proof of Concept By End of 2025, 2024. At least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025, due to poor data quality, inadequate risk controls, escalating costs or unclear business value. View source · verified 2026-06-20 · primary
  • S&P Global Market Intelligence — Generative AI shows rapid growth but yields mixed results (Voice of the Enterprise: AI & Machine Learning, Use Cases), 2025. The proportion of companies that abandon most of their AI initiatives has increased from 17% to 42%, with the average organization scrapping 46% of its proof-of-concept projects prior to production. View source · verified 2026-06-21 · ⚠ secondary mirror
  • EY — Get ready for robots: Why planning makes the difference between success and disappointment, 2016. as many as 30 to 50% of initial RPA projects fail. View source · verified 2026-06-20 · primary
  • Deloitte — The robots are ready. Are you? Untapped advantage in your digital workforce (Global RPA Survey), 2018. only 3% of organizations have managed to scale RPA to a level of 50 or more robots. View source · verified 2026-06-20 · ⚠ secondary mirror
  • DAIN Studios — How to Start and Scale Agentic AI: A.G.E.N.T. + 6-Step Strategy, 2026. Audit current workflows and roles, gauge value and complexity, engineer agent-first processes, navigate human-agent collaboration and track impact over time. View source · verified 2026-07-06 · primary
  • ServiceNow (community contributor) — The A.G.E.N.T. Method: A Practical Framework for Designing AI Agent Use Cases on ServiceNow, 2026. Before you open any tooling, understand the terrain. The framework recommends a 75-80% accuracy rate as the typical benchmark before advancing. View source · verified 2026-07-06 · primary
  • Harvard Data Science Initiative / Next Gen Learning — Agentic AI Intensive course page, 2026. A collaboration between Next Gen Learning and Harvard Data Science Initiative. View source · verified 2026-07-06 · primary