The method

A loop, a foundation, a thread.

Every cell runs the same five-stage loop, in cycles measured in weeks. The rails make it safe; the human thread makes the change stick.

The core idea

The future state is not known in advance — you discover what agents can reliably do by trying.

The shared, fatal assumption of the classic frameworks is that the destination is known in advance. With agentic AI you discover what agents can reliably do by trying — and the answer changes as the models improve.

Velora therefore does two things the classics do not: it redesigns the human role around the work agents take over, instead of deploying a tool and leaving the role untouched. And it never assumes a finished end state — the loop keeps turning as capabilities improve, so there is no “refreeze”.

The operating loop · five stages
1

Surface

Decompose roles into tasks; classify each task; map shadow-AI use; rank by value × feasibility × reversibility.

2

Pilot

Hand one task to an agent inside a small, reversible experiment. Design the human–agent handoff deliberately; instrument everything.

3

Evaluate

Calibrate trust with evidence. Measure value and reliability against the bar for the task’s risk tier; decide: scale, iterate, or stop.

4

Evolve

Redesign the role around the freed capacity. The stage the classics omit — and the one that decides whether adoption sticks.

5

Diffuse

Scale what worked, retire what did not, codify the pattern — and feed the loop back to Surface.

Exit criteria per stage. Surface: a prioritised shortlist of one to three tasks. Pilot: the workflow runs end-to-end with telemetry flowing. Evaluate: a documented go / no-go backed by evidence. Evolve: the role description and success measures reflect the new division of labour. Diffuse: the next team adopts the pattern faster than the first did.

What Surface must watch for

Boiling the ocean (trying to map everything) and mistaking a whole job for a single task. Resist staying at the level of job titles — decompose to discrete tasks.

What Pilot must watch for

Building a grand platform before proving one task; and piloting in a sandbox so unrealistic the results will not transfer.

What Evaluate must watch for

Vanity metrics (usage without value), trusting a good demo over measured reliability — and never killing anything.

What Evolve & Diffuse must watch for

The hollow reassurance; treating transition as training instead of redesign; refreezing (declaring victory and stopping); scaling by mandate instead of pull.

Governance by risk tier

Tiering is what lets you be fast and safe at the same time: full speed on low-risk work, real care on high-risk work. Set the tiers once, as a rail, and let every cell apply them.

TierTypical workHuman involvementExample
Tier 1 — low / reversibleInternal, low-stakes, easily undoneAgent runs autonomously; human spot-checksInternal drafts, first-pass data cleanup
Tier 2 — moderateCustomer-facing or money-adjacentAgent acts; human verifies before it commitsDraft client replies, prepared quotes
Tier 3 — high / regulatedLegal, financial, safety, or hard to reverseHuman-in-the-loop required, or task stays human for nowFinal approvals, regulated filings, hiring calls
The task-triage rubric · for Surface

Score each candidate task on three dimensions — value (how much is at stake), feasibility (how well an agent can do it today), and reversibility (how easily a mistake is undone) — and start where all three are high.

Delegate

Repetitive, rule-ish, low-risk, easy to verify. Pilot as Tier 1 — let the agent run, spot-check.

Co-do

Cognitively heavy but checkable; value in a human sign-off. Pilot as Tier 2 — agent drafts, human verifies.

Keep (for now)

Judgment, relationships, accountability, high irreversibility. Leave human; re-examine next loop as capability grows.

The human thread · executor to orchestrator

This is an identity change, not a skills upgrade — which is why it lives as a thread through every loop.

Borrowing from transition theory: name what is ending, support people through the uncertain middle, and make the new beginning concrete and desirable.

From — executorTo — orchestrator
Doing the workDirecting it
Producing outputVerifying it
Executing decisionsExercising judgment over them
Owning one taskOrchestrating several agents
The role-redesign playbook · for Evolve
  • List the tasks leaving the role — be specific about what the agent now does
  • Name what becomes more important — judgment, exceptions, oversight, relationships, directing agents
  • Write the new role on one page — new responsibilities, success measures, a new title if warranted
  • Have the honest conversation — acknowledge the ending, address the fear directly, describe the new beginning
  • Update the system around the person — performance metrics, incentives, and team structure, or the old role quietly returns
The cell

A cell is small — three to six people — and cross-functional. In a small organisation one person can wear two hats, but all four responsibilities must be present.

RoleWhat they own
Work ownerKnows the actual work and its quality bar; chooses tasks and judges outputs
BuilderConfigures the agent and the human–agent handoff
Risk ownerSets the tier, the guardrails, and what “good enough to trust” means
Sponsor (part-time)A leader who clears blockers, protects the cadence, and owns the honest role-change message
A 90-day first cycle
Outcome by ~day 90: evidence, redesigned roles, and a loop that keeps turning.
WeeksFocusKey outputs
1–2Stand up the railsDirection and honest role-change position stated; toolset sanctioned; risk tiers published; two to three cells chosen
3–4SurfaceTasks decomposed and ranked; shadow-AI use mapped
5–8Pilot + EvaluateOne human-plus-agent workflow per cell, measured against the tier bar
9–11EvolveAffected roles redesigned; transition conversations held
12–13Diffuse + re-SurfaceWinning patterns codified; failures retired cleanly; next backlog opened
Measuring it

Track a balanced set so the human side cannot silently fall behind the technical side. The trust metrics should fall over time as calibration improves.

What you trackExample metricWhat it tells you
Leading% of staff with sanctioned access; tasks surfaced; loop cycle timeWhether the machine is turning, and how fast
AdoptionActive usage; tasks delegated to agentsWhether capability is used, not just bought
ValueHuman time reclaimed; throughput; output qualityWhether the work is actually better, cheaper, or faster
TrustOverride rate; error / escape rate per taskWhether trust is correctly calibrated
PeopleRole-redesign completion; sentiment and safety signalsWhether the human side is keeping pace with the tech
Anti-patterns to avoid

Each common failure maps to a classic-framework assumption that no longer holds.

  • Pilot paralysis — endless experiments, nothing in production (no Diffuse stage, no cadence)
  • Suppressing shadow AI — banning what already works (ignoring grassroots adoption)
  • Tool rollout without work redesign — the most common cause of stalled adoption (skipping Evolve)
  • The hollow reassurance — promising no one will be replaced when it is untrue (it breaks the transition)
  • Designing for a final state — planning to refreeze (Lewin’s now-false assumption)
  • Big-bang transformation — one giant programme instead of many fast loops (Kotter at the wrong tempo)
  • Governance as afterthought — adding guardrails only after an incident

Start the loop in your organisation.

One cell, one pilot task, eight weeks — that is enough to have evidence and a redesigned role.