The tests are green. Why does the broken code still pass?
A Dart experiment shows why tests with full line coverage missed reversed pickup-point order, a radius boundary, and unavailable slots—and which assertions caught those faults.
- Published
- October 2, 2026
- Verified
- September 29, 2026
The tests are green. Why does the broken code still pass?
An app suggests pickup points for an order. The user expects the closest ones first: a short walk is preferable to a trip across town. The returned points are eligible, closed locations are filtered out, and the result respects the limit. The tests pass. Yet the list is now sorted in reverse, with the farthest point at the top.
That mismatch is easy to miss in review. There is a test for selecting pickup points; it calls the function and lists the expected IDs. The coverage report looks reassuring too. The useful question comes later: what failure did that test actually promise to catch?
ArkTelos Lab built a small Dart experiment using the ordinary package:test to examine that question. The pickup points are fictional, and the faults were introduced deliberately. No production bug is being claimed. The advantage of this setup is that the requirements, input data, source changes, and test results can be inspected together.
The right points, in the wrong order
selectPickupPoints receives a list that has already been prepared. Each record has an ID, a distance, an acceptsOrders flag, and a number of free slots. The function must retain points that accept orders, have at least one free slot, and fall within an inclusive distance limit. It then sorts them by increasing distance and returns no more than the requested number. Equal distances are resolved by ID. It does not change the input list.
That scope is intentional. Distances are already calculated, and availability is a single snapshot. There are no network requests or reservations inside this function. If availability becomes stale, correct sorting alone will not save an order. For now, the narrower question is whether a test can notice a breach of this contract.
Here is the relevant part of the function. The negative-argument checks that precede it are omitted for readability; they are present and tested in the saved source.
final selected = points
.where(
(point) =>
point.acceptsOrders &&
point.freeSlots > 0 &&
point.distanceMeters <= maxDistanceMeters,
)
.toList();
selected.sort((a, b) {
final distance = a.distanceMeters.compareTo(b.distanceMeters);
return distance != 0 ? distance : a.id.compareTo(b.id);
});
return selected.take(limit).toList();
Filtering, sorting, and limiting the result are three separate promises. A test can call the whole function yet check only one of them. That is what happened in the initial suite.
One of its tests is quite meaningful: given accepting, closed, and distant points, it expects the two eligible ones. Here are its inputs and main assertion together:
final points = [
PickupPoint('far', 900),
PickupPoint('closed', 300, acceptsOrders: false),
PickupPoint('near', 200),
PickupPoint('outside', 6000),
];
expect(
selectPickupPoints(points, maxDistanceMeters: 5000).map((p) => p.id),
unorderedEquals(['near', 'far']),
);
PickupPoint is an immutable record. Here, an unspecified freeSlots defaults to 1, and a point accepts orders by default. unorderedEquals checks membership without caring about order. That is exactly right for a result whose order has no meaning. But this function returns a ranked list, and that ranking affects the user's choice. After the distance comparator was reversed, the function returned ['far', 'near']. The test stayed green: both IDs were still there.
Another test requests two points but asserts only hasLength(2). Reverse a list of four eligible points and take the first two: the function now selects the two farthest, yet the length remains two. Each assertion answers its own question correctly. Neither asks whether the closest points come first.
The strengthened suite uses deliberately mixed distances: 900, 200, and 500 metres. The expected IDs come directly from the sorting requirement. In this second example, points is a new three-record list:
final points = [
PickupPoint('far', 900),
PickupPoint('near', 200),
PickupPoint('middle', 500),
];
expect(
selectPickupPoints(points, maxDistanceMeters: 5000).map((p) => p.id),
['near', 'middle', 'far'],
);
expect(
selectPickupPoints(points, maxDistanceMeters: 5000, limit: 2)
.map((p) => p.id),
['near', 'middle'],
);
All other conditions are satisfied for these three points. The expectation is written out rather than calculated by another sort inside the test. With the fault introduced, the full result became ['far', 'middle', 'near']; both new checks failed. The first caught the order, while the second also caught the wrong choice of the first two.
This does not make strict list comparison universally preferable. Where order is outside the contract, unorderedEquals is appropriate. The mistake is to discard a property of the result that the caller relies on.
Two more faults that passed unnoticed
The first gap was about order. The initial tests also checked distances inside and outside the allowed radius, but had no point exactly on its boundary. The condition distanceMeters <= maxDistanceMeters includes a point at 5,000 metres when the radius is 5,000. Changing <= to < leaves the ordinary examples green.
Adding points at 4999, 5000, and 5001 metres exposed the difference. The correct function returned ['inside', 'boundary']; after the change, only ['inside'] remained. The boundary case matters because the agreed wording is “no farther than,” which includes exactly 5,000 metres. The test needs an input on which those two interpretations diverge.
The other gap concerned free slots. The initial suite distinguished accepting points from closed ones, but every record had a positive freeSlots value. Removing freeSlots > 0 therefore changed nothing for those inputs. In the added example, both points accepted orders and were within the radius. The first had 0 free slots, the second had 1. The broken function returned ['full', 'available'] instead of ['available'].
All conditions on that input matter. Had full also been closed or too far away, another filter would still have removed it. Such a test could pass even after the free-slot check was deleted. To examine one rule, the other conditions must allow its violation to reach the output.
The initial tests were not window dressing. They caught two other deliberate faults: removal of the acceptsOrders check and returning one extra point beyond the limit. Empty input, a zero limit, and negative arguments were covered as well. The missing cases define the boundary of useful existing tests; they do not erase their value.
What the runs showed
Each change was applied independently to the same source file. Both suites first passed against the correct function. Then the unchanged tests were run for every source change. The function was restored and both suites passed again. The setup and results are saved with the experiment source.
| Deliberate behaviour change | Initial 6 tests | Strengthened 12 tests |
|---|---|---|
| Inclusive radius becomes strict | passed | detected |
| Farthest points come first | passed | detected |
| Free-slot check is removed | passed | detected |
| Closed points enter the result | detected | detected |
| One extra point is returned | detected | detected |
The initial suite detected 2 of the 5 selected behavioural faults; the strengthened suite detected 5 of 5. Those fractions describe only these five changes. They cannot estimate how many defects the tests would find in an application: the changes were hand-picked to explain differences between assertions, not sampled from production defects.
One detail makes the result less obvious. Both suites achieved the same line coverage of the function under study: 13 of 13 executable lines in the LCOV reports. That metric says which lines ran. It does not say whether the important combinations of input data were present or whether the output of every executed branch was checked against the requirement. The sorting line ran before the suite was strengthened. No assertion checked its order.
Coverage still answers a useful question: which code was never executed? A zero on an important line deserves investigation. But 100% line coverage does not turn hasLength(2) into a check for the nearest two points.
What if the source changes and tests stay green?
The technique now has a name: mutation testing. Introduce a small source change, then see whether the existing tests react. In this experiment, every change and both suites were fixed before the result matrix was inspected. The technique helps probe assertions, but a surviving change does not automatically mean the tests are weak.
One additional change replaced freeSlots > 0 with freeSlots >= 1. The field is an integer. These conditions are equivalent for integer values, so both suites passed as expected. Adding a test that must distinguish identical behaviour would be pointless. First ask whether any input in the declared domain can change the required result. Here none can. Such cases are called equivalent mutants.
The experiment also introduced an intentional syntax error. The analyzer reported it, but no tests executed. Recording that as “a test detected the defect” would give an assertion credit for a result it never checked. A runner needs to distinguish a failed named test from a suite that could not load.
The strengthened suite also checks ID ordering when distances tie and preservation of the input list. No separate source changes targeting those rules were tried. The “5 of 5” result therefore says nothing about how well those particular checks resist faults.
A question for test review
Instead of asking only whether the function is covered, take one obligation and imagine violating it. What will a user see if the farthest point comes first? What is the result exactly on the radius boundary? What happens when a point accepts orders but has no free slots? Then find the input and assertion that distinguish the correct answer from each violation.
This can be done in code review without running a dedicated tool. Take hasLength(2) and describe a broken result that still satisfies it: the function selects the two farthest points. The missing question is immediately visible. What to change depends on the cost of that mistake. For a ranked pickup list, checking the nearest two is useful. For a result with no specified order, replacing unorderedEquals with strict comparison would create a false requirement. A test should make the contract clearer, not accidentally narrow it.
This does not require one test for each line of a condition. A single scenario may check several obligations if its failures remain understandable. Conversely, ten checks around one happy path will not help with zero free slots until an input actually contains that value. A test count alone cannot tell us which decisions the suite protects.
There is a harder limit: the requirements themselves can be wrong. A strengthened test could faithfully encode an incorrect order if the team got “closest” wrong. This experiment began with distances already calculated. It did not test routing, availability freshness, reservation, the UI, or two people ordering concurrently. Those responsibilities need different scenarios and different checks.
The small ArkTelos Lab experiment reveals one specific confusion: executed code does not mean a checked obligation. A strong test fails when the function breaks the rule the test was written to protect. If that violation cannot be named, green only tells us that the examples used so far did not object.
More engineering articles are available at ArkTelos Lab. The ecosystem and its tools are on the official ArkTelos site.
Telegram: ArkTelos Lab EN | ArkTelos Lab RU | ArkTelos EN | ArkTelos RU.
Result boundaries
- Teaching Dart fixture with deliberately selected faults; 2/5 and 5/5 describe only these five. Coverage of 13/13 lines does not measure branch or scenario completeness or the correctness of requirements. UI, networking, availability freshness and reservations were not tested.
CODE / DATA / AGENTS
Related experiments
green-tests-mutation-trialCompare which deliberate breaches of a pickup-point selection contract are caught by initial and strengthened tests.
Teaching Dart fixture with deliberately selected faults; 2/5 and 5/5 describe only these five. Coverage of 13/13 lines does not measure branch or scenario completeness or the correctness of requirements. UI, networking, availability freshness and reservations were not tested.
Open the experiment record →Telegram edition