A convincing demo is not yet a client handoff. For a private OpenClaw support assistant, I would ask the builder to pass four tests before calling it deployed.
1. The input is limited to an approved channel and a named policy source.
2. A routine question produces a draft with its source; a missing or ambiguous policy produces an escalation, not a guess.
3. Refunds, account changes, promises, and uncertain answers cannot be completed without a human.
4. After a timeout or restart, an operator can see the request state, retry history, credentials used, and whether a reply was actually sent.
The handoff should include the knowledge boundary, escalation rules, and an operator runbook so someone else can recover the system. That is where the service value lives—not in a one-off answer.
The image sketches the automated-versus-human split; it is an illustration, not measured performance.
Which test would you insist on before deploying this for a client?