The agent has read the instructions. How can it inspect a running Flutter app?
The analyzer reports no issues and the model holds new orders, yet the screen keeps loading. An experiment examines what widget tests and MCP observations establish, and why an available tool does not necessarily become part of an agent’s verification.
- Published
- September 21, 2026
- Verified
- September 17, 2026
The agent has read the instructions. How can it inspect a running Flutter app?
An order list has a Refresh button. Press it, and a loading indicator appears. The source returns new orders and the operation completes. Yet the old list stays on screen, with the indicator still spinning.
The analyzer reports no issues. A test that inspects the model's fields passes too: the new orders are there, and the loading flag is false. An AI coding agent might identify the missing line immediately. The harder question is what makes the resulting patch acceptable: a plausible explanation, a passing test, or an observation from the application after the change?
Flutter agent skills and development tools address different parts of that question. Instructions can explain how a package should be used. A connection to a running application can expose what is happening now. Neither a list of installed skills nor a list of connected servers tells the reviewer which evidence the agent actually collected.
This ArkTelos Lab experiment follows a small UI defect through code, widget tests and a live runtime connection. It also asks whether an agent uses that connection when ordinary development tools remain available.
Instructions cannot report the current screen
A skill supplies instructions for a task: how to use an API, where to find its constraints, or what to check after a change. It can guide an investigation, but the instructions alone cannot report whether the loading indicator disappeared after the latest tap.
That requires an observation. The Dart and Flutter MCP server is one way to obtain it. MCP, the Model Context Protocol, provides a way for an AI client to call exposed tools. Here, those tools include Dart and Flutter diagnostics and access to a running application. The Flutter tooling guide distinguishes this local server from agent skills and documentation search.
In this article, the runtime is the application while it is executing, including its widget tree and available debugging information. Reading the button's implementation explains how it should work. Invoking the button in that process produces evidence about a particular run.
There is no reason to remove tests or shell commands to make MCP look necessary. Those tools remained available throughout the repair attempts. The question was what the agent would choose, not what it would do with every alternative taken away.
Correct fields, stale widgets
The fixture contains one screen, two orders, a Refresh button and a data-version label. Its source is local, with no network dependency. In the running app, the source responds after 300 milliseconds. Tests control the response explicitly.
The model stores a CatalogSnapshot, containing a version and an order list, alongside a refreshing flag and an error message. It extends ChangeNotifier, which lets it notify listeners of changes. A ListenableBuilder listens and rebuilds the relevant widgets.
The broken refresh method is short:
Future<void> refresh() async {
if (refreshing) return;
refreshing = true;
error = null;
notifyListeners();
try {
snapshot = await source.fetch();
} catch (_) {
error = 'Refresh failed. Previous orders retained.';
} finally {
refreshing = false;
}
}
The first notification tells the UI to display loading. Later, the model receives a new snapshot and clears the flag, but sends no completion notification. In the reproduced scenario, the widgets retain the state they built after the first notification.
A failed request has the same problem. The model has an error message and retains the old orders, but the UI has not been told to stop loading and display the failure.
Nothing here violates a type rule. Even a test can pass if it only asks whether the model received the new version. That test verifies data, not whether the data became visible.
The fixture deliberately avoids another state-management package or architectural layer. Neither is needed to expose this distinction. The broad error handler is also a simplification for the experiment, not a recommendation to handle every application error as the same string.
A test that checks the missing step
A widget test builds widgets, performs actions and inspects the resulting tree. For this defect, the useful question is not merely “did the model receive version 2?” It is “after the source responds, does the tree contain version 2 without the loading indicator?”
The test source implements CatalogSource, whose fetch() method returns a future snapshot. A Dart Completer lets the test decide when that future completes:
class ControlledSource implements CatalogSource {
final List<Completer<CatalogSnapshot>> requests = [];
@override
Future<CatalogSnapshot> fetch() {
final pending = Completer<CatalogSnapshot>();
requests.add(pending);
return pending.future;
}
}
After building CatalogApp with this source and its model, the test presses Refresh, checks loading, then supplies the response. This is the relevant excerpt:
await tester.tap(find.byKey(const ValueKey('refresh')));
await tester.pump();
expect(find.byKey(const ValueKey('loading')), findsOneWidget);
source.requests[0].complete(
const CatalogSnapshot(2, ['Order A / v2', 'Order B / v2']),
);
await tester.pump();
expect(model.snapshot.version, 2);
expect(find.text('Version 2'), findsOneWidget);
expect(find.byKey(const ValueKey('loading')), findsNothing);
pump() allows the test environment to process work and render a frame. Crucially, the test does not construct a fresh application after the response: doing so could hide the missing notification behind an unrelated rebuild. On the broken implementation, the model assertion passes, but the lookup for Version 2 fails.
The fix is one line:
} finally {
refreshing = false;
notifyListeners();
}
The notification follows the state change. A listener reading fields synchronously inside its callback must see the completed state, not refreshing == true.
The checks also cover a second refresh, a failed request that retains existing orders, and a successful retry. One successful tap would be too narrow a basis for accepting the whole behavior.
Observing the same behavior through MCP
The fixture also runs in flutter-tester, a headless Flutter runtime. It has a live widget tree without a separate graphical window. This is not an Android or iOS device test, and inspecting that tree does not establish pixel-level correctness.
The connection uses DTD, the Dart Tooling Daemon, which supports communication between Dart development tools. The MCP server receives the address for this particular application run. A test-only entry point enables Flutter Driver so the fixture's button can be invoked; this entry point is not intended for production.
The recorded environment uses dart_mcp_server 1.1.1, Flutter 3.44.7 and Dart 3.12.2. These are experiment versions, not a claim about the latest releases.
A separate script obtained the server's tool list, connected, read the tree, checked runtime errors and pressed Refresh. The difference was visible in the returned data:
| After the source response | Broken fixture | Fixed reference fixture |
|---|---|---|
| Version text | Version 1 |
Version 2 |
| Orders in the tree | Original v1 orders |
New v2 orders |
| Loading indicator | Still present | Absent |
| Runtime errors | None reported | None reported |
The last row matters. A program need not throw an exception to violate a user-facing requirement. “No runtime errors” does not establish that loading finished.
The fixed reference was started in a new process, rather than assuming an existing instance had adopted the patch. That is what a cold start means here.
Those observations came from a script, not from an AI agent. The next check therefore used the actual agent client: it had to discover the tools, obtain permission to call them and retrieve the result.
On the already-correct fixture, the agent received a short task: connect, read the tree, press Refresh and inspect the updated tree. The recorded tap used this tool name and argument object; this is not a terminal command:
{
"tool": "flutter_driver_command",
"arguments": {
"command": "tap",
"finderType": "ByValueKey",
"keyValueString": "refresh",
"keyValueType": "String",
"timeout": "3000"
}
}
The button has ValueKey('refresh'), so the call targets that widget. Before tapping, the check disabled Flutter Driver frame synchronization: a continuously animating loading indicator can otherwise interfere with interactions that wait for frames to settle. The agent then used widget_inspector with get_widget_tree and summaryOnly: true to read a compact tree, not a screenshot.
The first response contained Version 1 and the old orders. The second contained Version 2, the new orders and no CircularProgressIndicator. Runtime diagnostics reported no errors. All six MCP calls succeeded; the check took about 36 seconds.
That establishes the client's ability to perform this scenario. It does not establish the agent's ability to find a defect, apply a patch and verify its own change in one task: the application supplied for this check was already correct.
Available does not mean used
Three separate repair attempts followed, using GPT-6 Astra at medium reasoning effort through Codex CLI 0.153.4. Each received a clean broken fixture, the same task and an initial test that only checked whether the screen opened. The prepared fix, previous attempts and investigator-written evaluation tests were not supplied. Each attempt was limited to five minutes and 25 tool calls.
The task described the stuck list and required preserving loading, error recovery and repeat refresh. It included the running application's address, but did not require MCP verification as a condition of completion. The agent could choose how to check its work.
All three attempts found the missing notification and added regression tests. Those tests failed on the original code and passed after the fix. Each attempt took a little over two minutes and used 9–10 tool calls.
None of the three repair attempts called MCP. Each used code inspection, tests and the analyzer. Each final response also stated that the existing application instance had not been reloaded or checked. The agents did not claim observations they had not made.
Independent evaluation checked each patch with the analyzer and seven tests: the initial test, three participant-written tests and three investigator-written tests. All passed. A separate MCP script then verified refreshed orders and the absence of loading in a newly started application. That final observation belongs to the experiment's evaluation procedure, not to the agent that produced the patch.
This does not make MCP useless. The defect was small enough to identify in code, and its consequence was reproducible in a widget test. Runtime access was available, but was not necessary to produce these correct patches. Providing it did not make it part of the agents' verification.
An earlier six-attempt series had a client-approval configuration error that prevented MCP use. Its records are retained, but it is not treated as a comparison with working MCP. The successful client check and three new attempts followed the correction. Time limits and some task wording differ between the series, so their timings do not support a controlled performance comparison.
For the new check and three repair attempts together, the client reported 696,731 input tokens, including 588,288 cached input tokens, and 8,570 output tokens. These are cumulative usage figures, including context passed repeatedly, not the size of a single prompt or a monetary estimate. A model-based connectivity check consumes budget too.
What should count as completion?
For this screen, the requirement can be expressed without naming a tool: after a successful response, new orders and the new version are visible, loading has finished, and another refresh is possible. After failure, the old orders remain and a retry can succeed.
An agent can be asked to reproduce a violation, add a regression check, fix the code and rerun the check. If the task also requires observing the running application, that requirement should be explicit. For example:
After the fix, start the application with the changed code. Use MCP to press Refresh and check that the new orders and version appear and the indicator disappears. Report test results separately from runtime observations. If the connection or check is unavailable, state what remains unverified.
This is a proposed acceptance criterion, not a repair prompt validated by these trials. The short client check established that the actions can be performed on an already-correct fixture. How the agent would combine diagnosis, application restart and verification under this requirement remains untested.
“Use every connected tool” is a poorer instruction. It creates activity without specifying which uncertainty needs to be resolved. A regression test may be sufficient for a missing notification. A defect tied to a particular running state requires inspecting that state. A question about spacing, color or overlapping elements requires visual evidence, not just a widget tree.
Runtime access also needs boundaries. This experiment used only its own application address and withheld reference fixes and evaluation tests from participants. MCP project roots help tools navigate, but do not replace file and application access controls. Resistance to arbitrary malicious instructions was not tested.
The useful review question is therefore what behavior has been verified, and how. In this fixture, tests caught the defect and confirmed the repair. MCP separately exposed the same behavior in the running app. Simply adding the server did not cause the repair agents to perform that observation themselves.
ArkTelos Lab prepared the source, tests and experiment records for tracing the defect, the patches and the actual tool calls. To repeat the investigation, start by reproducing the stuck refresh and validating the connection. Set the model-run budget and completion criteria before adding agent trials.
ArkTelos: official website · news EN · новости RU
ArkTelos Lab: laboratory · Telegram EN · Telegram RU
Result boundaries
- Synthetic defect, Flutter 3.44.7 / Dart 3.12.2 / dart_mcp_server 1.1.1, headless macOS. One client probe on a correct app and three repair trials; participants did not use MCP. Cold checks were investigator-driven. No MCP superiority, pixel correctness, Android/iOS or prompt-injection security claim.
CODE / DATA / AGENTS
Related experiments
agent-runtime-verificationReproduce a missing notification and distinguish tests, MCP observations and agent actions.
Synthetic defect, Flutter 3.44.7 / Dart 3.12.2 / dart_mcp_server 1.1.1, headless macOS. One client probe on a correct app and three repair trials; participants did not use MCP. Cold checks were investigator-driven. No MCP superiority, pixel correctness, Android/iOS or prompt-injection security claim.
Open the experiment record →Telegram edition