Beyond the QA Bottleneck
How I used eight focused AI skills to bring tests, device evidence and review into React Native development, with monthly PR activity and production feedback.

In this article
AI adoption brought verification into the development loop
I made AI-assisted development include test authoring, device verification and evidence for review. From May through August 2026, the app recorded 556 merged PRs versus 161 in January through April, alongside a sustained build-out of the test suite. The earlier process had tests, but verification arrived too late: unit-test files moved from 93 in January to 113 in April, while Maestro stayed at 14 YAML files. Those device flows had existed since July 2024, yet they were not running in CI.
Finishing implementation left a familiar question: what else needed checking before we could release it? A nearby change could bring an already tested screen back into scope. Regressions weakened confidence, and manual QA carried repeated checks that did not leave an executable record for the next release. Useful backlog improvements became harder to schedule because the time between writing the code and trusting the result was difficult to predict.
I joined in April 2026 and started changing that process in May. The earlier chart history gives context for what followed. By 2 October, the app had 1,686 unit-test files and 232 Maestro YAML files, with smoke runs, device captures, review automation and installable test builds around the work.
My goal was to make AI adoption useful all the way through development. I wanted the agent to understand the requirement, implement it, write meaningful tests, run the checks and show the interaction on a device. That would give me something concrete to review. The change to the process mattered as much as the ability to generate code: each task needed to leave behind evidence and reusable protection.
We also moved towards product engineering. A smaller engineering team took on broader responsibility for product decisions, mobile development and the web experience. In practice, I saw features become faster to develop and deliver, with more work completed across that wider scope. That is my experience of the transition; the activity charts below provide a visible record alongside it.
The rhythm of development changed from May
The monthly record follows the same app repository from July 2024 through August 2026. Within that history, January through August 2026 gives a continuous comparison around the May change in process. Its merged-PR counts were 25, 32, 46, 58 from January to April, followed by 114, 151, 126, 165 from May to August. The app moved into a wider monorepo in September, so I keep later repository-wide activity out of this comparison. For longer-term context, the 2025 monthly counts ranged from 21 to 92.
The longer development record
Merged pull requests per month, July 2024 to August 2026
| Date | Merged pull requests per month, July 2024 to August 2026 |
|---|---|
| 2024-07-01 | 25 |
| 2024-08-01 | 26 |
| 2024-09-01 | 8 |
| 2024-10-01 | 51 |
| 2024-11-01 | 30 |
| 2024-12-01 | 38 |
| 2025-01-01 | 66 |
| 2025-02-01 | 92 |
| 2025-03-01 | 85 |
| 2025-04-01 | 69 |
| 2025-05-01 | 66 |
| 2025-06-01 | 38 |
| 2025-07-01 | 63 |
| 2025-08-01 | 35 |
| 2025-09-01 | 21 |
| 2025-10-01 | 92 |
| 2025-11-01 | 68 |
| 2025-12-01 | 40 |
| 2026-01-01 | 25 |
| 2026-02-01 | 32 |
| 2026-03-01 | 46 |
| 2026-04-01 | 58 |
| 2026-05-01 | 114 |
| 2026-06-01 | 151 |
| 2026-07-01 | 126 |
| 2026-08-01 | 165 |
The same application has a longer, variable history. The recent four-month comparison sits inside that record, rather than representing all earlier development.
Each point is a calendar month's merged-PR count, positioned at the start of that month. The line connects monthly observations. PR activity does not measure task size, productivity or customer value.
The marker shows the April observation, the month I joined. Workflow changes began in May.
Across the two four-month windows, that is 161 merged PRs before May and 556 from May through August. The monthly average increased from 40.25 to 139, about 3.45 times as many merged PRs. This records a sustained increase in merge activity. Changes differ in size and purpose, so the ratio cannot be read as the same increase in effort, productivity, features delivered or customer value.
Merged pull requests
Pull requests per month, 2026
- 25
- 32
- 46
- 58
- 114
- 151
- 126
- 165
Each bar starts at zero.
January to August 2026, one mobile application in the same repository.
Merged PRs measure activity, not productivity, effort, quality, distinct features or turnaround. September is excluded because the application moved and the later repository has a broader scope.
The marker shows the April observation, the month I joined. Workflow changes began in May.
I also count distinct ticket identifiers in each month. They moved from 22, 27, 33, 52 in January to April to 102, 139, 99, 130 in May to August. The sums are 134 and 470 monthly ticket occurrences. A ticket spanning two months can be counted in both, so I do not present those sums as unique features. This second measure helps check whether higher PR activity was accompanied by work against more tracked requests.
Tracked work
Distinct tickets per month, 2026
- 22
- 27
- 33
- 52
- 102
- 139
- 99
- 130
Each bar starts at zero.
January to August 2026, the same application. Tickets provide another view of work moving through the development process.
A ticket can cover a fix, maintenance or feature work, and can appear in more than one month. Monthly counts are not features shipped or a cross-period distinct total.
The marker shows the April observation, the month I joined. Workflow changes began in May.
PR counts are the clearest continuous record of development activity I can share here. I use them to show the rhythm of work moving through review, not as a productivity target. Ten PRs can involve more effort than fifteen, or less. Planning, investigation, implementation, testing and product decisions also happen without a merge, so a day with none can still contain substantial work. Splitting a feature into smaller PRs changes the count too. These monthly charts show recorded merge activity; they do not measure daily effort, feature completion or turnaround.
The outcome I cared about was how quickly we could develop and deliver useful features with confidence. I experienced an improvement there as the loop became established, even as the team became smaller and its product and web responsibilities grew. The PR history documents visible activity accompanying that change; it does not quantify the time from a feature request to its release or establish an improvement per engineer.
This is an observational comparison: team size, responsibilities, task mix, scope, AI assistance and the engineering process changed together. It does not isolate AI's contribution. The useful question is how we made the work assessable: which checks ran, what evidence accompanied it, and what remained uncertain at review.
I gave implementation and review the same rulebook
The first app rulebook appeared on 7 May. On 24 and 25 July, I made the standards explicit, moved lint rules from warnings to errors, connected the AI reviewers to those standards and added a self-verification runbook. The agent now had a contract to follow across different tasks.
The rules covered small files with focused responsibilities, business logic in hooks and helpers, reuse of existing components and tests for changed production code. The target was roughly 100 lines per file, rather than allowing one large UI file to accumulate every decision. Existing exceptions remained, so I treat this as an enforced direction for new work rather than claiming that the entire application suddenly met every standard.
The agent writes a plan in the task description before implementation. I can correct a misunderstanding while it is still a paragraph, before it becomes a feature to unwind. Focused skills then carry the approved task through the remaining work.
The same rules reach the reviewers. CodeRabbit reads the main guidance file; Qodo reads a mirror plus explicit compliance checks; cubic uses focused review rules. A small CI comparison fails when the mirror drifts from the main rulebook. Implementation and review use the same version of the contract.
Eight focused skills carry the task through the loop
I packaged the process as eight named skills with specific responsibilities. A skill is a repeatable set of instructions that an agent loads for a job. Cook coordinates the workflow, and saul keeps communication clear throughout. The other skills perform focused stages, each leaving a concrete output for the next one.
The 8 skills behind the workflow
cook
Orchestrate
Delegates each phase to a focused skill. It carries the task from discovery through the plan approval, isolated implementation, evidence and review, rather than asking one prompt to do everything.
Handoff: A task carried through a defined sequence.
sherlock
Discover
Detects the package manager, lint and format commands, types, build, test and coverage commands, CI names and PR conventions. It scopes checks to the repository and recognises work that already has a PR.
Handoff: The commands and conventions this task must satisfy.
dig
Research
Reads the requirement, design, discussion and code in parallel. Readers return relevant findings so planning can connect the sources without carrying every document into the implementation context.
Handoff: Findings grounded in the task's sources.
mastermind
Plan
Combines the research into the intended behaviour and a bounded checklist in the task description. It stops for my approval before implementation, where a misunderstanding is still cheap to correct.
Handoff: An approved checklist for implementation and verification.
michelin
Implement and verify
Implements to the written standards, adds tests and runs lint, types, build, tests and coverage. It diagnoses failures and repeats the checks, then assesses correctness and edge cases against the requirement.
Handoff: Changed code, meaningful tests and fresh check results.
mugshot
Capture
Re-examines the change, then exercises the affected interface in its actual screen. Separate workers claim separate iOS and Android devices, check feature flags and capture screenshots plus recordings for the task.
Handoff: Device evidence that a reviewer can inspect.
parole
Review
Opens the PR, keeps it current, evaluates bot and human feedback, fixes CI and resolves justified review threads. It stops at the human merge decision rather than accepting its own work.
Handoff: A current PR with checks and review concerns addressed.
saul
Communicate throughout
Shapes task updates, the PR body and review replies in my voice: short, direct, accurate and kind. It distinguishes a bot's suggestion from a person's question and explains the evidence behind a decision.
Handoff: Clear updates and review explanations throughout the task.
The handoffs keep autonomy bounded. Research returns findings from the ticket, discussions, documents, design and code; planning turns them into a checklist. Once I approve it, implementation works in a fresh branch and worktree from the main branch. Parallel tasks keep their checkouts and devices separate.
Capture and review have their own instructions and evidence requirements. Human review still controls acceptance and merge. A later run confirms the merge before clearing the temporary worktree, branch and local evidence and releasing the devices. Worktree creation and cleanup are lifecycle stages around the eight named skills.
This split also gives new lessons a home. Flaky selectors update capture guidance; a build command updates discovery; a missed edge case updates implementation and review. The next task gets those instructions alongside the tests that the previous work left behind.
A task had to produce more than implementation
The workflow became plan, implement, test, verify, capture, review and deliver a test build. Its completion criteria extend beyond the code: a changed behaviour needs tests, a UI change needs device evidence, and a reviewer needs fresh results for the revision they are assessing.
Verification travels with the change
Plan
Expected behaviour and a focused task
Implement and write tests
A change with repeatable checks
Verify
Test results and checks against the task
Capture device evidence
Screenshots and a recording of the flow
Review
A human decision supported by evidence
Deliver a test build
An installable build for real-device feedback
Observe and learn
Production signals that shape the next task
New learning returns to planning
The agent runs checks and reads their output before describing them as passing. Review material identifies what was exercised and its revision. If the branch changes, the affected checks need fresh results. Verification becomes part of finishing the task.
Device-evidence blocks started appearing in PRs on 19 June. Recent evidence tables link green device-flow runs to the revision under review. Those tables were still posted by the engineer, so I do not describe the whole evidence package as fully automatic.
This is the mechanism behind reducing repetitive QA. When a defect becomes a test, the next change can exercise that behaviour again. When a recording accompanies the implementation, a reviewer can inspect the result without reconstructing the whole task. When a test build is available, someone can try it on a real device. Each artifact answers a particular question and reduces how much confidence must be assembled again at release time.
The test inventory grew as tests became part of each change
The unit-test history shows the change in habit over several months. January to April ended at 93, 95, 98, 113 files. The following month-end snapshots were 164 in May, 250 in June, 395 in July and 1,040 in August. The app reached 1,364 files on 12 September and 1,686 on 2 October. Compared with April, that final snapshot contains 1,573 additional test files, about 14.9 times the earlier inventory.
Unit test files, January to October 2026
Files in the same mobile application
- 93
- 95
- 98
- 113
- 164
- 250
- 395
- 1,040
- 1,364
- 1,686
Each bar starts at zero.
The inventory grew from 93 files at the January checkpoint to 1,686 on 2 October. The April checkpoint had 113 files.
Bars show actual snapshots, including 12 September and 2 October, across a repository move. File counts are not individual tests, coverage or test quality.
The marker shows the April observation, the month I joined. Workflow changes began in May.
These are test files, not individual assertions or a coverage percentage. The 2 October inventory also contained 7,609 it() blocks and 225 parameterised it.each tables. A file can contain many cases, and splitting a file changes the inventory without increasing protection. I include the timeline because it makes the sustained investment visible; it needs to be read alongside the behaviours the tests exercise and the results of actually running them.
I made tests part of the implementation contract. The agent uses Jest with jest-expo and a shared render helper that supplies providers and routing. The test setup has 35 recorded MSW GraphQL handlers shared with 106 Storybook stories, so tests and visible states use the same controlled responses. Business decisions that do not need rendering live in pure helpers. That gives the agent a practical way to test a failure case without booting the entire interface.
For a regression, I ask for a test that fails before the fix and passes after it. I also temporarily revert the correction to confirm that the specification fails for the original reason. In one review, Qodo caught two defects that earlier passes had missed; both new specifications failed when their fixes were reverted. That establishes protection more clearly than a count or a test that follows the implementation.
Technique reference: React Native Testing Library
CI checks the change, and its cost became visible
The PR check chain includes Biome, TypeScript, unit tests, coverage, Expo export and prebuild, plus checks for the device flows and release tooling. AI agents write implementation and tests; these CI commands remain deterministic. The agent can run the chain, interpret a failure and correct the code, but its explanation cannot establish that a build passed. I want the actual output from the actual revision.
What happened across 100 app CI runs
Workflow runs, recorded 3 October 2026
- Successful
- 70
- Superseded
- 23
- Failed
- 7
The most recent 100 application CI runs at the snapshot. Superseded runs are a separate outcome from failures.
These are application CI workflow outcomes, not native-device results or public releases. The recorded median workflow duration was 15.6 minutes.
Coverage needs an explicit scope. File counts describe an inventory, while a coverage report describes code exercised during a particular test run. Expectations for changed production code and inherited project-wide checks concern different scopes. A green global result cannot establish complete application coverage or prove that every changed line met the task's expectations. I review the changed behaviour and its relevant test results together.
Consolidating the app and shared packages made more context available to agents and removed a separate publishing cycle for a component change. I observed CI time move from about five to fifteen minutes; the last 100 app CI runs recorded on 3 October had a median of 15.6 minutes. The snapshot classified 70 as successful, 23 as superseded and 7 as failed. A superseded run is neither a failure nor a completed successful verification. Longer CI was a real cost to the feedback loop, making affected-only checks the next prerequisite.
I also separated dispatch success from test success. GitHub Actions triggers the native work in EAS; a dispatch job can finish in seconds while the build and flows continue elsewhere. Fresh scheduled device builds tie the run to a known commit. Their results need to be visible where someone makes the merge or release decision.
Maestro became a running system rather than a folder of flows
The Maestro inventory stayed at 14 YAML files from January through April, grew to 129 in May, then reached 135 in June, 136 in July and 147 in August. It stood at 162 on 12 September and 232 on 2 October. The checkpoints show investment in reusable device checks over time; 232 files include several roles rather than 232 independent journeys.
UI flow files, January to October 2026
Maestro YAML files in the same application
- 14
- 14
- 14
- 14
- 129
- 135
- 136
- 147
- 162
- 232
Each bar starts at zero.
The first four checkpoints held at 14 files. The May snapshot had 129, and the October snapshot had 232.
Bars show actual checkpoints. This inventory includes leaf flows, shared subflows, suite aggregates and other YAML files. It is not a count of distinct journeys, runnable tests or coverage on every platform.
The marker shows the April observation, the month I joined. Workflow changes began in May.
The first EAS device workflow landed in May. On 18 June, the flows moved into Maestro Cloud with PR smoke and nightly execution. That was the important change from the earlier dormant folder: a test could now run alongside development and return a result. Making that useful required deliberate suite entrypoints, controlled setup and jobs that could finish within the device service's limits.
Technique reference: Maestro: how device flows work
Leaf flows, shared setup and aggregates have different jobs
The snapshot records 125 leaf flows, 53 shared sub-flows and 37 suite aggregates within the 232-file Maestro inventory. Leaves exercise a behaviour, shared sub-flows reuse setup or common steps, and aggregates provide entrypoints that call the relevant leaves. Those role counts describe the structure; they are not an exhaustive, mutually exclusive partition of every YAML file or a count of independent journeys.
Three roles inside the device-test inventory
Recorded YAML file-role counts
- Leaf flows
- 125
- Shared subflows
- 53
- Suite aggregates
- 37
The application inventory contained 232 YAML files. These three recorded roles explain how individual checks, shared setup and suite entrypoints work together.
The role counts are not an exhaustive, mutually exclusive partition of all YAML files. They do not measure independent journeys or coverage.
Selection follows the call graph. Each cloud job names an aggregate entrypoint; tags on individual files do not decide which CI runs include them. A new leaf therefore needs to be connected to the entrypoint that should exercise it. Shared setup reduces repeated maintenance, but changing it can affect several callers. The static analyser checks paths and reachability so that adding a plausible file does not silently leave it outside execution.
Device execution needed platform scope and a time budget
The scheduled configuration reached 31 cloud jobs: 22 on iOS and 9 on Android. These are configured suite jobs, rather than 31 passing tests or a coverage percentage. PR labels select a smoke run or the broader device suite. Fresh scheduled builds tie execution to a known commit, and each platform's results describe the work it actually ran.
Nightly device jobs by platform
Configured Maestro Cloud jobs
- iOS
- 22
- Android
- 9
The scheduled configuration contained 31 jobs, each naming a suite entrypoint for a platform.
A configured job can call several flow files. These counts describe execution structure, not passing tests, distinct journeys or equal platform coverage.
We changed the structure after a flaky suite prevented the suites behind it from producing evidence. On 5 August, nightly execution was split into one job per suite and platform. Two days later, a nine-leaf run hit Maestro Cloud's 15-minute limit. Some leaves needed roughly two minutes of setup each, so we split long suites into parts of three leaves or fewer. A flow that fits the execution budget can finish and leave a useful result; grouping everything into one long run made failures harder to isolate.
Predictable setup mattered as much as grouping. The flows use controlled initial state and launch arguments to select test conditions, with backend scenarios for known responses rather than a separate mock server. Accessibility labels and exact text matching make selectors easier to diagnose, while overlays and sticky elements can still intercept taps. I want a failed flow explained and repaired. Changing test data until a run happens to pass leaves the original uncertainty in place.
I added static checks for the tests themselves
Growing the device suite exposed another problem: a file could look plausible and still be unable to run, or never be reached by a scheduled aggregate. On 5 August, I added a static analyser for the Maestro YAML. It checks syntax, paths, application identifiers, selectors that cannot resolve and which flow files the CI entrypoints can reach. That catches problems before an expensive build and device run.
The analyser has nine rules, eight of them blocking, and 104 self-test cases. It runs in under a second in the normal CI lint job. Exceptions live in an allow-list with reasons rather than being silently ignored. The distinction between an inventory file and a reachable, executable test became part of the tooling, rather than something a reviewer had to remember every time.
We documented failure patterns for agents: exact text matching, footers intercepting taps, overlays and incorrect initial state. The next agent reads those learnings before trying the same approach. Debugging updates the test and its maintenance instructions.
Release automation has its own self-tests, with 167 checks in the PR chain. Tags decide versions, and over-the-air updates must reject native changes. Those rules need executable specifications because a regression in the delivery machinery can affect the result too.
The review includes screenshots and a recording from both platforms
Device evidence runs as a second pass after implementation and checks. The agent launches the app, reaches the affected screen and exercises the change where it is actually used. If a flag gates the interface, it confirms the flag is enabled before recording. An isolated component example helps with states, but the evidence also needs to show the surrounding screen and interaction.
The workflow asks for multiple screenshots and an MP4 on each platform. One reviewed change attached twelve files: three screenshots and three recordings on iOS, and the same on Android. The iOS capture uses simulator recording converted to MP4; Android uses screenrecord and pulls the recording with adb. Screenshots show visible states, while recordings show waiting, failure or recovery. Platform and revision details travel with the captures so that a reviewer can understand what they represent.
iOS and Android capture can proceed in parallel with one worker per platform, provided each claims a separate device and pins its commands to that device identifier. This prevents two concurrent tasks from tapping or recording the same simulator. Recordings are attached to the task, and the PR links to that evidence. It gives the reviewer a direct path from the requested behaviour to what happened on the device.
Simulator captures show the interaction in that environment; TestFlight adds feedback from physical devices. Reviewers can inspect the recorded result and decide what exploratory checking remains, instead of beginning acceptance by asking the author to recreate the changed flow.
AI review became another pass over a concrete change
From 25 July, three AI reviewers read the development rules alongside human reviewers. In the recorded sample of the last 30 merged PRs, they wrote 360 of 635 inline comments, about 57%. The remaining 275 came from other reviewers. This sample covers the broader repository, including work beyond the app. It measures review activity, not comment correctness, bugs prevented or the share of implementation AI wrote.
AI reviewers added another review pass
Inline comments across 30 merged PRs
- AI reviewers
- 360
- Other reviewers
- 275
AI reviewers wrote 360 of the sample's 635 inline comments, about 57%. This sample includes work beyond the mobile application.
Comment counts measure review activity. They do not establish comment correctness, bugs prevented or the share of code written by AI.
The implementation agent evaluates each comment: inspect the code, reproduce the concern, change when justified and rerun the affected checks. A reviewer can reveal a missing failure case or misunderstand a deliberate choice. Replies need evidence; automatically accepting suggestions is not a quality gate.
The PR brings the requirement, implementation, test output, device captures and remaining limitations together. Human review controls acceptance and merge, with a concrete decision to make and a visible place to record it. Test builds extend assessment beyond the author, and TestFlight gives testers access on physical devices.
I established dedicated test-build workflows on 22 June. Their purpose is to prepare installable builds for feedback beyond local captures, using distribution appropriate to each platform. A completed submission gives testers another way to assess the change on a real device. Publishing a public store release remains a separate decision with its own steps.
Technique reference: Expo: distributing a build with TestFlight
A more active release history needed clear delivery milestones
The recorded release history contains eight entries between 24 August and 30 September, roughly five and a half weeks. That shows a much more active release line, but the evidence does not enumerate eight confirmed public-store go-live dates on both platforms. A tag, a production workflow run, a TestFlight submission and a public release are different events. I keep those distinctions visible rather than using the eight entries to claim multiple public releases every week.
On 11 and 12 September, I added release tooling and runbooks to make delivery more repeatable. Generated release notes derive from reviewed PR titles, and self-tests check the version and delivery decisions against the workflow configuration. The release machinery needs executable specifications just as the application does.
I separated the delivery milestones: creating a build, submitting it for testers, preparing a store release and making it public. Automation can support these steps, but success has to be established at the step itself. A successful dispatch is not confirmation that the downstream build or submission completed. A release-history entry therefore needs context before it becomes evidence of public availability.
Native store builds and over-the-air updates need different compatibility checks. I added tests around the release rules, including rejecting native changes from an update intended for an existing runtime. Human review and public-release decisions remain part of the process. The improvement was a repeatable delivery path and a clearer account of what each recorded event established.
The Sentry history became part of the development loop
On 28 July, I made error cleanup systematic and connected Sentry issues to development tickets. A resurfacing problem could return to the backlog instead of being rediscovered in an unrelated chat. Release context and active feature flags accompany events, helping explain which implementation the user encountered. The work then becomes reproduce, classify, fix where appropriate and add a test that preserves the correction.
Daily recorded errors, July to October 2026
Events per day, all environments
| Date | Events per day, all environments |
|---|---|
| 2026-07-07 | 24,012 |
| 2026-07-08 | 20,186 |
| 2026-07-09 | 18,996 |
| 2026-07-10 | 19,756 |
| 2026-07-11 | 23,024 |
| 2026-07-12 | 23,245 |
| 2026-07-13 | 16,295 |
| 2026-07-14 | 15,949 |
| 2026-07-15 | 16,897 |
| 2026-07-16 | 15,846 |
| 2026-07-17 | 17,738 |
| 2026-07-18 | 16,197 |
| 2026-07-19 | 17,393 |
| 2026-07-20 | 10,078 |
| 2026-07-21 | 7,617 |
| 2026-07-22 | 9,141 |
| 2026-07-23 | 6,265 |
| 2026-07-24 | 9,794 |
| 2026-07-25 | 8,768 |
| 2026-07-26 | 10,718 |
| 2026-07-27 | 8,498 |
| 2026-07-28 | 8,371 |
| 2026-07-29 | 5,913 |
| 2026-07-30 | 5,496 |
| 2026-07-31 | 5,007 |
| 2026-08-01 | 4,590 |
| 2026-08-02 | 5,165 |
| 2026-08-03 | 4,218 |
| 2026-08-04 | 4,651 |
| 2026-08-05 | 4,593 |
| 2026-08-06 | 4,182 |
| 2026-08-07 | 3,842 |
| 2026-08-08 | 4,780 |
| 2026-08-09 | 4,634 |
| 2026-08-10 | 4,177 |
| 2026-08-11 | 3,653 |
| 2026-08-12 | 4,288 |
| 2026-08-13 | 4,385 |
| 2026-08-14 | 5,033 |
| 2026-08-15 | 5,656 |
| 2026-08-16 | 4,126 |
| 2026-08-17 | 3,445 |
| 2026-08-18 | 4,120 |
| 2026-08-19 | 4,262 |
| 2026-08-20 | 4,618 |
| 2026-08-21 | 5,010 |
| 2026-08-22 | 4,141 |
| 2026-08-23 | 5,675 |
| 2026-08-24 | 3,831 |
| 2026-08-25 | 4,252 |
| 2026-08-26 | 4,041 |
| 2026-08-27 | 5,122 |
| 2026-08-28 | 4,364 |
| 2026-08-29 | 5,519 |
| 2026-08-30 | 4,933 |
| 2026-08-31 | 4,120 |
| 2026-09-01 | 5,575 |
| 2026-09-02 | 2,856 |
| 2026-09-03 | 3,542 |
| 2026-09-04 | 4,435 |
| 2026-09-05 | 3,839 |
| 2026-09-06 | 4,728 |
| 2026-09-07 | 3,747 |
| 2026-09-08 | 4,707 |
| 2026-09-09 | 4,024 |
| 2026-09-10 | 3,937 |
| 2026-09-11 | 5,324 |
| 2026-09-12 | 5,981 |
| 2026-09-13 | 6,254 |
| 2026-09-14 | 6,468 |
| 2026-09-15 | 6,111 |
| 2026-09-16 | 5,340 |
| 2026-09-17 | 4,087 |
| 2026-09-18 | 7,077 |
| 2026-09-19 | 5,791 |
| 2026-09-20 | 8,333 |
| 2026-09-21 | 4,433 |
| 2026-09-22 | 8,013 |
| 2026-09-23 | 4,019 |
| 2026-09-24 | 4,336 |
| 2026-09-25 | 5,115 |
| 2026-09-26 | 5,466 |
| 2026-09-27 | 5,764 |
| 2026-09-28 | 5,162 |
| 2026-09-29 | 4,156 |
| 2026-09-30 | 4,189 |
| 2026-10-01 | 3,847 |
| 2026-10-02 | 5,941 |
| 2026-10-03 | 4,155 |
- 6 July to 4 August, 202630 dates; first hour missing
- 378,937 recorded events
- 4 September to 3 October, 202630 complete days
- 154,779 recorded events
Complete days from 7 July to 3 October. The history shows the fall and the daily variation that a two-total comparison would hide.
All environments are included. Fixes, filters and sampling affect recorded volume. Counts are not adjusted for sessions or usage and are not a crash rate or proof that AI caused the change. The incomplete 6 July and 4 October days are excluded from the line.
The daily chart shows all-environment error events from the retained history. Its completed days run from 7 July through 3 October. The 6 July boundary was missing its first hour, and 4 October was still in progress when the snapshot was taken. I exclude those incomplete endpoints from a completed-day trend. A very low count from a partially elapsed current day would otherwise make the apparent improvement look much larger than the evidence supports.
A reconstructed 6 July to 4 August window contains 378,937 events, with that first hour missing. The 30 complete days from 4 September to 3 October contain 154,779, about 59% fewer recorded events. Fixes, reporting filters and sampling changes all contributed; the data cannot separate their effects. All environments are included, and there is no session or traffic denominator. These are recorded event counts, not a crash-free rate or a measure of bugs eliminated by AI.
Recorded volume rose again during September before settling, something a two-point chart would hide. Distinct issues and repeated events tell different stories: one recurring condition can dominate the count. I use these signals to prioritise investigations. Reproducible defects return to the same test, device-evidence and review loop.
Technique reference: Sentry: release health session statistics
Performance needed its own evidence and controls
On 26 August, I added a performance harness that samples main-thread activity during repeated list flings on a Release simulator build. It performs ten 300-millisecond flings, so repeated runs have a defined workload. One finding was that Session Replay accounted for 48% to 73% of UI-update work in that measurement. That figure describes the sampled workload; it is not a claim that disabling Replay made the whole application that much faster.
I introduced performance changes behind feature flags so they could be measured before temporary controls were retired. The source records 20 performance flags retired, including 14 on 22 September. The useful habit was to profile a known workload, make a bounded change and rerun the measurement before simplifying the implementation.
I also added static checks for patterns that had already created avoidable work: timers, expensive blur and gradient effects, autoplay without pause handling, large shadows and unstable style objects. The linter blocks new occurrences of eight selected patterns while documenting older exceptions. This turns a profiling lesson into a check that the next agent encounters before it repeats the same approach.
A local profile and a scheduled performance trend establish different things. The local finding describes a known workload and environment. Using scheduled measurements as a trend requires completed, repeatable runs with results that can be compared. Performance belongs in QA only when its workload, environment and results are visible enough to interpret; configuring a workflow alone does not establish protection for every release.
A useful loop also reports what it did not verify
The 232 Maestro files do not mean every path runs on every platform. They include shared setup and suite entrypoints, while the configured jobs cover different platform scopes: 22 on iOS and nine on Android. A green result needs to say which behaviours it exercised. Inventory growth and a passing subset cannot establish complete platform coverage.
Review needs completed device and performance results, rather than confirmation that a job was triggered. I want the result linked to its revision, platform and verification scope, where a reviewer can inspect it. If that evidence is absent, the review should say what remains unverified instead of inferring success from the workflow configuration.
The next improvements follow those scopes: broaden platform verification, keep test setup predictable, make device and performance results easy to inspect, and verify delivery paths from the reviewed change through their downstream result. Affected-only CI matters too, because the move from roughly five to fifteen minutes makes repeated verification more expensive. Each improvement needs a defined scope and an observable result.
What I take from these months is a practical way to adopt AI: define the work, give the agent clear rules, make it write and run meaningful tests, capture the interaction and put the evidence in front of a reviewer. The measured result was a more active app development record and a much larger verification inventory. In practice, feature development and delivery became faster while a smaller team took on broader product and web responsibilities. The lasting benefit is that the next task starts with tests, shared components and production learning that the previous work left behind.
Resources for putting this into practice
These official references explain the tools and measurement concepts. The observations in the figures are my anonymised aggregate records, not results reported by these documentation sources.
- Maestro: how device flows work
Accessibility-driven interactions and reusable UI flows.
- React Native Testing Library
Tests that exercise components through user-facing behaviour.
- Expo: distributing a build with TestFlight
Build submission, tester access and the public-release boundary.
- Sentry: release health session statistics
Session and user counts, crash rates and crash-free rates.
