AI proposed a change. What did the user actually approve?
An assistant's proposal does not authorize a task update. A Dart experiment examines confirmation against displayed data, stale revisions, repeated requests and recorded operation outcomes.
- Published
- September 11, 2026
- Verified
- September 6, 2026
AI proposed a change. What did the user actually approve?
A planning assistant suggests moving a task's due date from September 10 to September 12. The application shows the suggestion, the user confirms it, and the new date is saved. That makes a convincing integration demo: the model understood the request, returned structured data, and the application acted on it.
Then someone edits the task while the confirmation screen is still open. Or the change succeeds, but the caller receives no response and tries again. Or the handler takes the date from the model's latest message while the screen still shows an earlier suggestion.
All three flows have a confirmation button. The button does not establish which change was authorized, or whether that change is still valid.
This boundary is worth examining before choosing another agent SDK. It determines what the application must save when preparing an action, what it must check before execution, and how it should respond to a retried request. Leaving these decisions inside a button handler makes application behavior depend on the screen or delivery mechanism that happens to invoke it.
The ArkTelos Lab: Agent action approval experiment explores that boundary in Dart. Fixed inputs stand in for model responses. The subject is the application's behavior after receiving a suggestion, not generation quality, network transport, or the model's ability to choose a sensible date.
A suggestion becomes an application-owned proposal
The input contains three fields:
{
"action": "reschedule_task",
"taskId": "task-1",
"newDueDate": "2026-09-12"
}
Before presenting anything, the application validates the action, field types,
and date. 2026-02-30 is not a valid date merely because it has the right
shape. Extra fields are rejected too: a model-supplied approved field must
not quietly acquire the authority to permit a write.
Even valid input cannot establish whose authority the operation uses, whether the task exists, or whether that user may change it. The application obtains those facts from its own caller context and current data.
Once those checks pass, the application saves a proposal: a record of one
specific change. The code calls it Proposal. It contains the task ID, the
user who initiated the proposal, the old and new dates, the original data
version, and an expiry time. The application also assigns the proposal's own ID.
The task version is an integer, revision, incremented whenever the task
changes. The proposal captures it as expectedRevision. Execution is allowed
only while the task still has the version on which the proposal was based.
The screen needs to display this saved record. Taking the displayed due date from one message and the due date applied by the handler from another breaks that relationship; matching task IDs cannot repair it.
In the experiment, proposals live in memory and cannot be modified through the returned collection. Preparing one leaves the task's due date unchanged. The application has recorded a possible operation, not performed it.
Approval must identify the change
A flag such as approved = true seems sufficient while there is only one
possible change. Once arguments can change, it becomes ambiguous: which
proposal was approved, with which date, and against which original data?
Here, confirmation carries the proposal ID and displayedIntent, a structured
set of the values presented for approval. The handler compares it with the
saved proposal, checking every field and value type, including the original
revision and expiry. Changing the proposed date produces
confirmation_mismatch. It does not damage the original proposal, which can
still be confirmed with matching data.
Sending every value back is one implementation choice. A system can instead submit only the ID of an immutable record, provided it maintains a reliable link between that record and the user's decision. The requirement is the link, not a particular confirmation payload.
Matching values also do not prove that a person saw the screen and clicked the button. A model tool allowed to invoke confirmation as the user could simply copy the expected values.
The example therefore exposes two separate interfaces. The suggestion adapter
receives a ModelPort with prepare; the user-decision adapter receives a
UserPort with confirm and decline. Both pass the user's identity separately
from the model's JSON. These interfaces are small enough to inspect in the
handler implementation.
This separates dependencies; it is not a security sandbox for hostile code inside the process. A real application must authenticate user requests and control which operations its model tools can invoke. The experiment uses scripted user decisions, not a real confirmation UI.
The task can change while approval is pending
Consider a task due on September 10 at revision 7. A proposal would move it to September 12. Before confirmation, another edit moves the date to September 15 and advances the revision to 8.
Saving September 12 without another check would overwrite an edit the user might never have seen. The user approved a change from September 10; the application would actually change the date from September 15.
In scenario A10, confirmation returns stale. The task stays at September 15,
revision 8. Continuing requires a new proposal based on current data and a
new user decision.
Checking the whole revision is deliberately conservative. It can reject a proposal after an unrelated field changes. That is the cost of requiring the original snapshot to remain current. Checking only the date, or merging independent changes, is possible, but needs a more precise definition of acceptable conflicts.
Permissions can change as well. Access granted during preparation may have
been revoked by the time the user confirms. Approval does not restore it.
In A09, the handler returns permission_revoked without changing the task.
The example also gives proposals a five-minute lifetime. This is a fixture
setting, not a universal recommendation. A14 checks two separate cases:
execution is permitted one microsecond before expiry and rejected with
expired exactly at the boundary.
These checks govern a new execution. Retrieving an outcome already recorded requires a different order.
A repeated request need not repeat the change
Suppose the new date is saved but the caller receives no answer. Disabling a button cannot settle this problem: the same confirmation might arrive from another screen or be retried by a delivery mechanism.
The application needs a record of how that proposal ended. On success, the
experiment stores a receipt containing the proposal ID, old and new dates,
before-and-after revisions, and commit time.
The proposal ID also identifies its execution. If a final outcome already
exists, an exact confirmation retry returns that outcome with replayed = true.
It does not enter the task-update branch again. This is the handler's limited
idempotency guarantee: repeating the same request does not apply the change
again.
After checking the proposal's author, current read access, and matching
confirmation data, confirm looks for that saved outcome first:
if (record.outcome != null)
return Reply.result(record.outcome!, replayed: true);
if (!allowWrite) return _finish(record, 'permission_revoked');
if (!now.isBefore(DateTime.parse(intent['expiresAt'] as String)))
return _finish(record, 'expired');
Here, record holds the proposal and its outcome, intent contains the saved
arguments, and _finish records a terminal rejection. The checks preceding
this excerpt remain part of the full method.
Expiry prevents a new execution under an old proposal. It does not undo a completed change or turn a request for its outcome into an expiry error. In A18, the first return value is discarded and the clock advances by a day. Retrying returns the original outcome; the date remains September 12 at revision 8.
A19 then considers a later edit. After successful execution, the task moves to September 20 at revision 9. Retrying the old confirmation still returns the receipt for September 10 to September 12. The task's current due date is untouched.
That distinction matters to the UI. A historical receipt establishes that an operation happened. The task's current state must be read separately. Using an old response to replace current screen data remains dangerous even when duplicate execution has been prevented.
The change and its outcome need a single commit
If the application updates the task first and records success second, a failure between those steps leaves a changed task with no execution record. A retry cannot find the record it relies on. Reversing the order creates the opposite problem: a success record exists before the task has changed.
The experiment holds tasks and proposal outcomes in one immutable Snapshot.
The handler prepares a new snapshot containing both the changed task and its
successful outcome:
final next = Snapshot(
{..._root.tasks, task.id: changed},
{..._root.proposals, id: Record(record.proposal, outcome)},
);
if (failBeforeCommit) return _finish(record, 'failed_before_commit');
_root = next;
return Reply.result(outcome);
_root is the current snapshot, changed is the new task version, and
outcome is the success result. Until _root = next, the previous data remains
unchanged. There is no asynchronous wait or external call between the final
check and that assignment.
The failBeforeCommit flag introduces a controlled failure after constructing
the snapshot but before adopting it. In A15, the date stays at September 10,
revision 7, with no success receipt. A terminal failed_before_commit outcome
is recorded instead. A16 repeats the confirmation and receives that same
outcome. A fresh attempt needs a new proposal and approval.
That policy distinguishes repeated delivery from a user's new attempt. It does not mean every application must ask again after every technical failure.
The single-commit property depends on the actual execution boundary: one process, one Dart isolate—an isolated execution context—and synchronous in-memory operations. All records disappear when the process ends. A database or an action in an external service needs its own definition of how the change and execution record are committed. Moving this code into an HTTP handler does not provide that guarantee.
Connecting the operation to Flutter
A screen needs a pending proposal, the confirmation outcome, and perhaps a notification to display. They serve different purposes: the proposal describes a currently available change; the outcome records how the operation ended; the notification asks the UI to perform a presentation action.
Several presentation designs can preserve these distinctions. ark_mvp, for example, separates the application snapshot, screen-ready state, and presentation effects. ark_mvp_flutter connects that boundary to Flutter's widget tree. These references describe the 1.0.0 contracts. The distinction between state, results, and effects is explored further in the earlier ArkTelos Lab article.
The application handler still owns approval rules and protection against repeated execution. Presentation packages do not provide authorization or a store of completed commands. A success notification cannot replace the saved outcome either.
No Flutter integration was built for this experiment. These responsibilities help frame such an integration; they do not establish how a screen behaves when closed or restored.
What the experiment actually exercised
Local verification used Dart 3.12.2 on macOS ARM64. The analyzer reported no issues, and 24 tests passed: scenarios A01–A23 plus a snapshot-immutability test. The CLI ran the same 23 scenarios. That is another way to inspect their results, not 23 additional independent tests.
| Situation | Observed outcome |
|---|---|
| Proposal prepared, no confirmation | Task unchanged |
| Confirmation arguments altered | Rejected; original proposal remains intact |
| Task revision changes before confirmation | Old proposal rejected; newer task state preserved |
| Identical confirmation repeated | Previous outcome returned; revision does not advance |
| Task edited again after success | Old receipt returned without overwriting current data |
| Controlled failure before commit | Task unchanged; no success receipt |
Inputs, assertions, and observations are available in the scenario definitions, JSON results, and verification record. A lost response is simulated by discarding a return value. Actual network disconnection and process crashes were not tested.
The repository instructions explain how to reproduce the run. Use the linked commit when comparing results. The local verification script creates a disposable copy with a separate dependency cache, then removes both. No model calls or paid API are needed.
The useful starting point for an application review is now concrete. Trace where proposal data comes from, which change the user approves, what is checked immediately before execution, and where a retry's answer comes from. If any transition depends only on the model's latest message or the button's state, that is where the application contract needs attention.
The next experiment worth running would add persistent storage and a process failure between request and response. It would test which properties survive beyond one application's memory. This experiment supplies specific expected outcomes for that work.
More engineering investigations: ArkTelos Lab. Related packages and documentation: the ArkTelos website.
Channels: ArkTelos Lab EN | ArkTelos Lab RU | ArkTelos EN | ArkTelos RU.
Result boundaries
- Fixed inputs and scripted user decisions; no live model or real confirmation screen was tested.
- One process, one isolate, synchronous memory. No network, durable storage, process-recovery or multi-process concurrency testing.
- ark_mvp and ark_mvp_flutter are discussed as a presentation boundary; Flutter integration is outside the experiment.
CODE / DATA / AGENTS
Related experiments
agent-action-approvalExamine confirmation against a saved proposal, current permissions and task revision, and retrieval of a recorded outcome on retry.
Fixed inputs and scripted user decisions; no live model or real confirmation screen was tested. One process, one isolate, synchronous memory. No network, durable storage, process-recovery or multi-process concurrency testing. ark_mvp and ark_mvp_flutter are discussed as a presentation boundary; Flutter integration is outside the experiment.
Open the experiment record →Telegram edition