The method
A loop, a foundation, a thread.
Every cell runs the same five-stage loop, in cycles measured in weeks. The rails make it safe; the human thread makes the change stick.
The future state is not known in advance — you discover what agents can reliably do by trying.
The shared, fatal assumption of the classic frameworks is that the destination is known in advance. With agentic AI you discover what agents can reliably do by trying — and the answer changes as the models improve.
Velora therefore does two things the classics do not: it redesigns the human role around the work agents take over, instead of deploying a tool and leaving the role untouched. And it never assumes a finished end state — the loop keeps turning as capabilities improve, so there is no “refreeze”.
Surface
Decompose roles into tasks; classify each task; map shadow-AI use; rank by value × feasibility × reversibility.
Pilot
Hand one task to an agent inside a small, reversible experiment. Design the human–agent handoff deliberately; instrument everything.
Evaluate
Calibrate trust with evidence. Measure value and reliability against the bar for the task’s risk tier; decide: scale, iterate, or stop.
Evolve
Redesign the role around the freed capacity. The stage the classics omit — and the one that decides whether adoption sticks.
Diffuse
Scale what worked, retire what did not, codify the pattern — and feed the loop back to Surface.
What Surface must watch for
Boiling the ocean (trying to map everything) and mistaking a whole job for a single task. Resist staying at the level of job titles — decompose to discrete tasks.
What Pilot must watch for
Building a grand platform before proving one task; and piloting in a sandbox so unrealistic the results will not transfer.
What Evaluate must watch for
Vanity metrics (usage without value), trusting a good demo over measured reliability — and never killing anything.
What Evolve & Diffuse must watch for
The hollow reassurance; treating transition as training instead of redesign; refreezing (declaring victory and stopping); scaling by mandate instead of pull.
Tiering is what lets you be fast and safe at the same time: full speed on low-risk work, real care on high-risk work. Set the tiers once, as a rail, and let every cell apply them.
| Tier | Typical work | Human involvement | Example |
|---|---|---|---|
| Tier 1 — low / reversible | Internal, low-stakes, easily undone | Agent runs autonomously; human spot-checks | Internal drafts, first-pass data cleanup |
| Tier 2 — moderate | Customer-facing or money-adjacent | Agent acts; human verifies before it commits | Draft client replies, prepared quotes |
| Tier 3 — high / regulated | Legal, financial, safety, or hard to reverse | Human-in-the-loop required, or task stays human for now | Final approvals, regulated filings, hiring calls |
Score each candidate task on three dimensions — value (how much is at stake), feasibility (how well an agent can do it today), and reversibility (how easily a mistake is undone) — and start where all three are high.
Repetitive, rule-ish, low-risk, easy to verify. Pilot as Tier 1 — let the agent run, spot-check.
Cognitively heavy but checkable; value in a human sign-off. Pilot as Tier 2 — agent drafts, human verifies.
Judgment, relationships, accountability, high irreversibility. Leave human; re-examine next loop as capability grows.
This is an identity change, not a skills upgrade — which is why it lives as a thread through every loop.
Borrowing from transition theory: name what is ending, support people through the uncertain middle, and make the new beginning concrete and desirable.
| From — executor | To — orchestrator |
|---|---|
| Doing the work | Directing it |
| Producing output | Verifying it |
| Executing decisions | Exercising judgment over them |
| Owning one task | Orchestrating several agents |
- List the tasks leaving the role — be specific about what the agent now does
- Name what becomes more important — judgment, exceptions, oversight, relationships, directing agents
- Write the new role on one page — new responsibilities, success measures, a new title if warranted
- Have the honest conversation — acknowledge the ending, address the fear directly, describe the new beginning
- Update the system around the person — performance metrics, incentives, and team structure, or the old role quietly returns
A cell is small — three to six people — and cross-functional. In a small organisation one person can wear two hats, but all four responsibilities must be present.
| Role | What they own |
|---|---|
| Work owner | Knows the actual work and its quality bar; chooses tasks and judges outputs |
| Builder | Configures the agent and the human–agent handoff |
| Risk owner | Sets the tier, the guardrails, and what “good enough to trust” means |
| Sponsor (part-time) | A leader who clears blockers, protects the cadence, and owns the honest role-change message |
| Weeks | Focus | Key outputs |
|---|---|---|
| 1–2 | Stand up the rails | Direction and honest role-change position stated; toolset sanctioned; risk tiers published; two to three cells chosen |
| 3–4 | Surface | Tasks decomposed and ranked; shadow-AI use mapped |
| 5–8 | Pilot + Evaluate | One human-plus-agent workflow per cell, measured against the tier bar |
| 9–11 | Evolve | Affected roles redesigned; transition conversations held |
| 12–13 | Diffuse + re-Surface | Winning patterns codified; failures retired cleanly; next backlog opened |
Track a balanced set so the human side cannot silently fall behind the technical side. The trust metrics should fall over time as calibration improves.
| What you track | Example metric | What it tells you |
|---|---|---|
| Leading | % of staff with sanctioned access; tasks surfaced; loop cycle time | Whether the machine is turning, and how fast |
| Adoption | Active usage; tasks delegated to agents | Whether capability is used, not just bought |
| Value | Human time reclaimed; throughput; output quality | Whether the work is actually better, cheaper, or faster |
| Trust | Override rate; error / escape rate per task | Whether trust is correctly calibrated |
| People | Role-redesign completion; sentiment and safety signals | Whether the human side is keeping pace with the tech |
Each common failure maps to a classic-framework assumption that no longer holds.
- Pilot paralysis — endless experiments, nothing in production (no Diffuse stage, no cadence)
- Suppressing shadow AI — banning what already works (ignoring grassroots adoption)
- Tool rollout without work redesign — the most common cause of stalled adoption (skipping Evolve)
- The hollow reassurance — promising no one will be replaced when it is untrue (it breaks the transition)
- Designing for a final state — planning to refreeze (Lewin’s now-false assumption)
- Big-bang transformation — one giant programme instead of many fast loops (Kotter at the wrong tempo)
- Governance as afterthought — adding guardrails only after an incident
Start the loop in your organisation.
One cell, one pilot task, eight weeks — that is enough to have evidence and a redesigned role.