ATL-2026-011benchmarkexperimental

JSON parsing moved to an isolate. Which problem did that solve?

Offloading JSON parsing can delay the result while reducing animation stalls. A Flutter experiment on M1 shows why the optimization criterion must be defined before choosing an API.

Published
September 25, 2026
Verified
September 23, 2026

JSON parsing moved to an isolate. Which problem did that solve?

A loading indicator freezes while an application processes a large server response. The developer moves parsing into a separate isolate, and the animation stops freezing. Then a timing measurement reveals that the result now arrives later.

It is easy to call the change a failure: the operation became slower. It is equally easy to defend it: the indicator is moving again, so the application has been optimized. Both assessments leave something important out. The first ignores what happened to the interface during processing. The second says nothing about how long the user now waits for the data.

That is how a careful benchmark can justify reverting a useful fix. Or how an unhelpful change can survive because it looks convincing on screen.

Before choosing an execution strategy, there is a less technical question to answer: which behaviour needs fixing? The wait for a result, an animation stall, or the inability to interact during loading? These requirements can concern the same operation, but need different checks. The word “performance” makes it too easy to pretend that one number describes all of them.

Finishing sooner is not the only requirement

In Flutter, synchronous JSON parsing and application-object construction can occupy the same main isolate that needs to service the interface. While that work runs, the isolate cannot move on to the next event in its event loop. Adding async to the surrounding method does not move the work elsewhere.

Concurrent execution needs another context with its own event loop: a separate isolate. It receives a task, and the caller gets the result later. That separation has a cost: startup, messaging and object delivery. It does not promise a shorter wait.

A concrete result from the ArkTelos Lab experiment makes the distinction visible. The same prepared string was converted into a complete object list in three ways: on the main isolate, in a fresh isolate for each request, and in a persistent isolate. The last one is called a worker below. In this fixture it handles requests sequentially and remains available after each one finishes.

On a MacBook Air M1, a Flutter application in profile mode processed JSON containing 100,000 records, approximately 14.6 MB. A small square rotated on screen; after processing, only the record count changed. Each strategy handled 50 requests across five ten-second windows.

Execution strategy Median time until the full list was ready Median longest frame-callback gap around a request
Main isolate 210.1 ms 219.2 ms
Fresh isolate 273.1 ms 20.4 ms
Persistent worker 339.4 ms 23.2 ms

For the second numeric column, each request contributes its longest interval between frame callbacks, measured from processing start until completion plus two frame budgets. A frame callback is how the application observes a frame; the interval is not an exact count of frames missed by the display. The screen ran at 60 Hz, giving an update budget of approximately 16.67 ms.

The fresh isolate returned the list about 63 ms later, but reduced the characteristic long gap from 219 to 20 ms. If the goal was to reduce animation stalls during parsing, that is a useful change with a measured cost. If the only requirement was to obtain the complete list sooner, the main isolate remained ahead on that criterion.

This does not establish that users always prefer a longer load with a moving indicator. The experiment tested neither user preferences nor button-response latency. It showed something narrower: data completion time and the ability to service frames moved in opposite directions in this scenario. The product requirement must determine whether that trade-off is acceptable.

If the screen disables every action throughout loading anyway, moving the work may improve only the indicator. That still provides useful feedback, but preserving interaction has not yet been demonstrated. That requires testing the available actions, not just watching an animation.

A convincing metric can answer a different question

The same main-isolate run offers another number: the p95 of all frame-callback intervals was approximately 17.2 ms. This is the threshold within which at least 95% of the recorded intervals fall. That summary gives little reason to suspect recurring pauses longer than 200 ms.

No calculation error is needed. A request ran once per second, leaving many ordinary frames between requests. The relatively rare long gaps fell beyond the chosen percentile. The p95 of frame build duration was only 0.47 ms: short work after a pause does not describe the wait before it starts.

The metric changed meaning when it was separated from the user action. “Most frames fit within the budget” and “processing a response does not cause a long pause” are different statements. The first can be true while the second is false.

Checking a fix therefore requires looking at the interval in which the complaint occurs. If the screen freezes after receiving a response, inspect both that processing and the frames around it. Recording more uneventful seconds must not turn the same freeze into an apparently successful fix.

This is not a reason to replace every statistic with a maximum. A single maximum may reflect an unrelated event: the no-JSON control recorded a 55.8 ms frame-callback interval. Repetition, association with the operation and a control are needed. An aggregate remains useful as long as it is clear which observation it preserves and which it hides.

Where should the offloaded work end?

Once the decision to offload is made, a function name can become a tempting boundary: the expensive jsonDecode runs elsewhere, while everything else stays as it was. But a screen rarely needs only the output of syntactic parsing. It needs objects it can use.

Model construction, field validation, filtering and sorting may sit between those points. Moving only the first step leaves the remaining synchronous work on the main isolate. Its cost needs measuring; the presence of a separate isolate does not settle that question.

The fixture places the boundary at the completed result. parseCatalog parses the string and eagerly creates the entire List<CatalogItem>, including nested tag lists. A CatalogItem is a record with an identifier, a name, an integer price, an availability flag and tags. There is no lazy conversion left to construct records after the return.

Future<List<CatalogItem>> parseFresh(String input) =>
    Isolate.run(() => parseCatalog(input), debugName: 'catalog-fresh');

This excerpt comes from the fixture; the class and parser definitions are available in the full source. The helper sits outside the screen object and receives only the required input. That matters: a closure can implicitly capture extra state, and exchanged objects must be sendable between isolates. See the Isolate.run documentation.

A particular screen may need a different completed result, such as an already filtered and sorted selection. Preparing that selection can be considered one stage, rather than returning intermediate data merely because jsonDecode has finished.

There is no need to move the entire screen object along with the data. Presentation stays on its own side; the computation receives the information it needs. A useful boundary depends on what the next stage requires and who can prepare it without knowing about widgets.

Filtering was not measured in this experiment. It is a subsequent design decision, not another improvement supposedly demonstrated by the table. Even after complete objects arrive, building a heavy interface can introduce its own pause: data processing does not eliminate presentation work.

What does a persistent worker actually preserve?

A persistent worker looks like a natural next optimization. Instead of creating an isolate for every request, start it once. Yet in the table this strategy delivered the large list later than the others. “Startup is expensive, so reuse is cheaper” was not enough to explain the complete operation.

The return mechanisms differ. Isolate.run terminates its isolate and transfers the result through Isolate.exit without copying. A worker that remains alive sends a message through SendPort.send: immutable objects such as strings may be shared, while other parts of the structure are copied. A large list of instances with nested lists makes that work relevant. See Isolate.run and SendPort.send.

The measurements cannot attribute the entire gap to copying: copying, garbage collection and thread scheduling were not isolated. But the difference between the mechanisms already exposes what the startup argument leaves out. Keeping a processor alive and delivering its result cheaply are separate tasks.

A more interesting question comes earlier: why does the caller need the complete catalogue every time?

Suppose a screen displays a small selection while the remaining records support later searches. The catalogue could stay beside the worker, which receives search criteria and returns only selected records. This changes the amount of data exchanged; the fixture did not measure whether it improves timing.

Such a worker takes on a new responsibility: storing data and performing operations on it. The design must specify when the catalogue refreshes, which version an answer represents, who owns changes and when storage is disposed of. If both sides modify data without a clear rule, lower transfer costs can introduce a separate consistency problem.

For a one-off load, that may be unnecessary complexity. For repeated queries over one large structure, it may be a sensible architecture. The choice depends on who needs the data and how long it lives. The number of lines in a worker implementation does not answer either question.

This is also where an impressive but unfair benchmark becomes easy to construct. The main isolate returns the complete catalogue, while the worker returns a count or a small selection. The second operation transfers less data and fulfils a different contract. The designs can be compared against the screen's actual task, but that is not evidence that the original operation became faster while producing the same output.

Moving a computation does not decide whether its result is still needed

Once the main isolate is free, the application may accept new actions while an earlier computation continues. That improvement has consequences beyond the parsing-time table.

A user might change search criteria before the previous request finishes. Its result remains correct for the old query but no longer fits the current screen. Unconditionally updating presentation after every await can therefore leave a more responsive interface displaying stale data.

Applying a result needs a rule: verify that it belongs to the current request, or associate it with an input version. Rejecting an outdated response does not save computation; the work still runs. Cancellation and queue limits require separate decisions.

Request races and cancellation were not investigated here: the fixture deliberately permits only one active worker task. Its results do not establish that the interface can now handle a new request on every keystroke.

A useful review outcome is therefore a clear description of behaviour: which work stops interfering with the screen, what comes back when it finishes, and under which conditions that result is still needed. Milliseconds verify part of that description. The rest does not arrive automatically with Isolate.run.

For the heavy request in this experiment, a fresh isolate was sufficient to substantially reduce animation stalls. A persistent worker would need a separate justification tied to reuse or data ownership. With the small, 100-record input, none of the strategies showed a comparable frame problem. There, the reason to introduce another execution context still needs establishing.

A frozen-indicator fix cannot be judged solely by how soon it returns a list. A slow-loading fix cannot be considered complete merely because the indicator rotates. And if a screen needs only a small selection, the need to move an entire catalogue across the execution boundary deserves examination first. These are three different reasons to change the code. While they remain bundled into “speed up JSON”, even the right API can solve the wrong problem.


The experiment ran on a MacBook Air M1 with 16 GB RAM, macOS 26.5.1, Flutter 3.44.7 / Dart 3.12.2, on battery. The table above covers Flutter profile mode only. A separate ahead-of-time compiled Dart series, other input sizes, warmup details and the full protocol are in the experiment materials. Original results are preserved, including two excluded attempts with the window hidden; observations were not removed because their values were large. An ordinary desktop environment is not a controlled laboratory. These measurements establish neither Android, iOS or Web behaviour, nor memory use, energy consumption or input latency. On the Web, compute runs its callback on the current event loop, so the native conclusions cannot be transferred there. See the compute documentation.


ArkTelos · The lab
Telegram: Lab RU · Lab EN · ArkTelos RU · ArkTelos EN

Result boundaries

  • Local experiment on MacBook Air M1, 16 GB, macOS 26.5.1, Flutter 3.44.7 / Dart 3.12.2, on battery. The table covers Flutter profile, with 50 requests per strategy and size. No Android/iOS/Web, memory, energy, input latency, user preference, cancellation or request-race measurements. Copying cost was not isolated. Two hidden-window attempts were excluded in full and preserved. Publication preparation did not rerun the experiment.

CODE / DATA / AGENTS

Related experiments

experimentjson-isolates-trial

Compare full-list completion time and frame-callback gaps across three JSON execution strategies.

Local experiment on MacBook Air M1, 16 GB, macOS 26.5.1, Flutter 3.44.7 / Dart 3.12.2, on battery. The table covers Flutter profile, with 50 requests per strategy and size. No Android/iOS/Web, memory, energy, input latency, user preference, cancellation or request-race measurements. Copying cost was not isolated. Two hidden-window attempts were excluded in full and preserved. Publication preparation did not rerun the experiment.

Open the experiment record →

Telegram edition

Read in ArkTelos Lab

Read in ArkTelos Lab ↗

ARKTELOS CHANNELS

News and engineering material, without mixing languages.

The main channels cover ArkTelos development. ArkTelos Lab publishes architecture analysis, experiments, and reproducible research.