CI and performance metrics
How Matrix engineers measure test feedback, profile slow suites, and validate CI improvements.
This guide is for Matrix engineers and contributors who maintain CI. For local test commands and test conventions, use the testing guide.
Metrics and the five-minute target
The five-minute CI target is an acceptance goal, not a measured guarantee. Report queue and execution separately, and retain the full time from a commit to the final required check.
| Metric | Definition | What it tells us |
|---|---|---|
| Queue time | Time from a job becoming eligible to a runner starting it | Runner availability; dependent jobs are not eligible until their prerequisites pass |
| Execution time | Time from runner start to job completion, including setup and cleanup | Cost of the job itself |
| End-to-end time | Time from the triggering commit to completion of all required checks | The feedback delay a contributor actually experiences |
| Setup time | Checkout, dependency installation, browsers, and prerequisite builds | Work that can benefit from reuse or caching |
| Test time | Wall-clock time for the test command | Execution cost after setup; summed file durations can overlap under parallelism |
| Coverage and outcomes | Passed, failed, skipped, and expected test counts | Whether the faster run still validates the same behavior |
Compare the same suite scope and required checks. Keep the commit, runner shape, worker count, image digest, cache state, counts, and failure status with each measurement. Report cold and warm passes separately. An incomplete, partial, interrupted, or failed run is diagnostic evidence; it does not establish a complete CI speedup. A snapshot of fixture startup is not a whole-suite benchmark.
Current rollout status
The work beginning with Matrix OS PR #2442 remains an unmerged CI stack. The latest complete delegated Linux qualification passed on 11 October 2026 in 11 minutes 02 seconds, with no queue wait. It includes the newly landed native Paste changes and the optimization stack. Required hosted CI remains the merge gate during rollout. Priority routing has not been activated, and the five-minute target remains unmet.
Latest qualification and diagnostic evidence
The original full GitHub CI run passed all 17 jobs in 65 minutes 43 seconds, including 20 minutes 22 seconds waiting on its critical path. Its type-check command took 72 s within a 144 s job, and its Web build command took 128 s within a 189 s job. Command timings exclude setup and remaining job overhead.
| Measurement | Elapsed | Outcome and scope |
|---|---|---|
| Original full hosted CI | 65 min 43 s | All 17 jobs passed; includes hosted gates and queue wait |
| Gallery fix main CI | 53 min 04 s | All 17 jobs passed; approximately 9 min 27 s waiting on the critical path |
| Native Paste main CI | 55 min 50 s | Required CI Results gate passed; 16 jobs passed, one conditional Docs job skipped; approximately 10 min 25 s waiting |
| Earlier optimized Linux qualification | 9 min 26 s | Previous source; full delegated workload passed |
| Latest optimized Linux qualification | 11 min 02 s | Updated source; full delegated workload passed; queue wait zero |
These are different source revisions and timing boundaries. The latest Linux qualification does not measure the complete hybrid GitHub workflow, including retained hosted gates and controller overhead. The difference from hosted CI is therefore an indication of the opportunity, not a measured whole-CI speedup. Release publication is a separate workflow and is excluded.
The latest unit report contains 27,927 passed, zero failed and 274 existing skipped test cases: 28,201 total across 2,688 test-file reports. General browser checks passed 194 cases with 104 existing skips; the terminal-grid checks passed seven cases; Electron Desktop checks passed 48 cases with two macOS-only skips. All 55 measured phases and 14 source-integrity guards passed, including the new native Paste assertions. Container removal and evidence validation completed successfully.
| Final measured phase | Elapsed |
|---|---|
| Setup, including four frozen dependency installations | 1 min 45 s |
| Full unit lane, 16 workers | 9 min 10 s |
| Native type checks, eight projects | 17 s |
| Web production build | 3 min 17 s |
| Electron Desktop production build | 15 s |
| General browser regressions | 3 min 23 s |
| Terminal-grid browser checks | 19 s |
| Required Electron regressions, nine separate groups | 3 min 10 s |
| Sync package build / tests / publish check | 13 s / 16 s / 1 s |
| SDK install / runtime compatibility | 4 s / 10 s |
| Final validation, collection and cleanup after the critical lane | 7 s |
Independent lanes overlap, so their durations cannot be added to obtain total elapsed time. The unit lane remains the critical path; the complete browser lane took 7 min 08 s. A faster Web build alone will not meet the five-minute target. Sharding can improve balance, but adding more workers to one already busy host can increase contention. Multiple PRs will also wait for the single full-workload slot.
The qualification source was 5677467e345b32a459fac625f971549ef7496ac0: the optimization stack combined with the public source tree of Native Paste PR #2454, which landed as f54f8414f875e4125e92aba92762a40951a8ea27. It is not the subsequently rebased rollout source. The run used a Helsinki Ryzen 9 7950X3D host with 16 physical cores, 32 logical threads and 128 GB nominal RAM. Its disposable sandbox used 30 CPU threads, a 112 GB memory limit and 16 unit workers. The immutable image digest was:
sha256:f1f759e6493097476d0110ac40bd7a6010afa38f64c0f57b14a68f2923bf010bThe earlier 9 min 26 s qualified source (1c3ccfaa10e7a46a20a7ef7383e148233c04581d) had 26,853 unit passes and 272 existing skips. Its smaller inventory cannot be compared as an isolated compiler or hardware experiment. Failed or incomplete historical candidates remain diagnostic evidence only.
Node 20 compatibility, Pattern Scan, React Doctor, funded PostgreSQL and privileged host contracts remain hosted checks. Their queue and complete GitHub aggregation have not been measured as part of this Linux qualification. Routing stays disabled until current-source review, full qualification and ordinary and stacked-PR shadow canaries pass.
Native TypeScript checks
The standalone checker pins native TypeScript 7.0.2 for all eight no-emit project checks. Both native and legacy checkers passed those scopes. JavaScript TypeScript 5.9.3 remains available for declaration builds and framework compiler/API integrations; this is not a blanket compiler replacement. The checker preserves the previous ambient-type and side-effect-import scope, and real gateway inference errors were fixed rather than suppressed.
Prepared caches and authenticated routing
The prepared runner image retains immutable dependency and browser bytes. A submitted lockfile must match the prepared lock before reuse; each disposable run receives its own writable package store. Independent lanes can run concurrently, while source integrity and missing-output checks still invalidate stale build caches.
The proposed priority workflow requires both ci-linux and ready-for-ci on an open, non-draft, same-repository PR. Engineers should wait for current-head review before admitting it. Labeling a PR will not provide Linux routing until the repository switch is enabled.
One full Linux workload runs at a time. Other admitted PRs wait in order; their queue time is reported separately. A new commit supersedes only that PR's old work. A parent-only change in a stack invalidates the child's old source result and requires fresh checks against the new merge. A queued or completed result for an old head, parent, requesting attempt, image or harness cannot establish a passing current gate.
The trusted default-branch controller revalidates the exact current source after acquiring the server lock. Bounded leases stop an owned workload on cancellation, lost connection, expiry or source drift; cleanup must finish before the next workload starts. Candidate code receives no host credentials or Docker socket.
Shadow canaries keep the normal hosted lanes. The rollout must demonstrate ordinary PRs, non-main stacked bases, new-head supersession, parent-only changes, cancellation, queue preservation and rollback before replacing eligible heavy lanes. Main pushes, forks, unsupported events and hosted-only platform/security checks retain hosted coverage. Missing, stale, wrong-source, failed or incompletely cleaned-up dedicated evidence must fail the required merge gate.
For rollback, disable MATRIX_CI_DEDICATED_ENABLED and the shadow switch, remove admission labels as needed, and start fresh hosted checks for affected PRs. Disabling routing alone does not replace a run whose hosted lanes were already skipped. Confirm the current required hosted gate is green before merging.
Web build imports
The production trace now visits 216 Hugeicons icon modules, compared with 6,025 in the earlier trace. All 264 named icon exports retained their original object identity and rendered SVG markup. The scoped Next import optimization preserves the existing aliases, rewrites and release wrapper; canonical production builds and browser regressions passed. Module counts demonstrate reduced traversal, while elapsed build time still depends on concurrent load.
Electron Desktop build bridge
A standalone Linux experiment reduced Electron Desktop build time from 53 seconds to 7 seconds with exact rolldown-vite@7.3.1 and explicit Electron externals in main and preload. The unchanged Electron, clipboard, terminal-grid and security lanes passed 116 tests, with two macOS-only skips. Integrated build/security contracts and all eight native type checks also passed.
This package is a deprecated temporary migration bridge, not a maintained security minor. Stable Electron-Vite 5.0.0 currently accepts Vite 5–7; its Vite 8 successor is still beta. Vite's migration guide documents the intermediate path. Follow up with compatible stable Electron-Vite and Vite 8 releases, repeating runtime/security and macOS packaging checks before removing the alias and scoped override. The seven-day release-age policy remains enforced. These standalone timings do not certify whole-CI speed or macOS packaging/signing.
Profiling slow suites
Measure runner queue time, dependency setup, prerequisite builds, and test execution separately. A faster test process does not remove a slow runner queue.
The CI work starting with Matrix OS PR #2442 adds unit JSON profiles, with duration-balanced shards in a dependent layer. After those layers land, a local profile can be collected with:
bun run test:prepare
pnpm exec vitest run --reporter=json --outputFile=/tmp/matrix-unit.json
node scripts/ci/test-profile.mjs --root "$PWD" \
--output /tmp/matrix-test-durations.json /tmp/matrix-unit.jsonReview successful reports before updating the checked-in timing manifest. New test files stay included and receive an estimated duration until measured; balancing must never reduce coverage. CI worker overrides are bounded to 1–16, while ordinary local runs keep their default behavior.
Dedicated Linux benchmarks use an exact commit, a disposable container, and cold/warm passes. Keep credentials outside test code. Compare baseline and optimized commits on the same image and hardware, retaining test counts, timings, and failure status. The dedicated benchmark is a subset of validation: required hosted CI remains the merge gate until equivalent coverage is proven. A five-minute CI target is an acceptance goal, not a measured guarantee.
The proposed lane-parity correction keeps general E2E and required Electron regressions aligned with their hosted lanes. Building Electron before a general E2E run can activate additional native suites; those extra suites must be reported separately when comparing performance.
Generated E2E evidence can also invalidate a clean-checkout provenance check. The proposed evidence correction ignores six known generated directories while keeping unrelated untracked source visible. Provenance checks must still reject source changes; do not ignore the entire output tree to make a benchmark pass.
See the repository's benchmark protocol and operator guide for the proposed rollout. The operator guide supports 8-vCPU/32-GB and 16-vCPU/64-GB hosts. Automatic dispatch remains disabled pending validation and review.
Fresh database fixtures
The latest native PostgreSQL fixture migration covers 76 test files and 1,302 test cases. Those cases passed on both native and fallback paths with unchanged application assertions, followed by the passing full Linux unit qualification above. Module-owned closed templates produce independent clones, with bounded connections and explicit teardown. Native PostgreSQL is selected only by the fixture environment; migration, fresh-schema, shared-instance and timer-sensitive tests retain their existing setup unless separately qualified. This validated work remains unmerged.
Nine additional existing suites now use cached empty initialized engines instead of repeating engine initialization. Every suite still runs its own schema creation, migration and bootstrap, with independent data and cleanup. The count is nine migrated files, not a cache-cap setting. The earlier snapshot layers below describe the fallback path; they do not cache a live shared database.
The proposed platform fixture optimization retains one immutable, pre-migrated schema image per isolated Vitest module. Every fixture restores a separate engine and connection, then runs the normal startup revision checks. Tests must keep independent data, sequences, schema changes, and transactions; caching a live database would break that isolation.
Migration contracts opt out with createTestPlatformDb({ freshSchema: true }) to retain fresh initialization. Validate fixture contracts and existing database-heavy suites before adopting snapshot reuse. Local initialization timings do not establish complete CI performance.
The proposed Chat fixture helper uses an empty initialized PostgreSQL image in the Chat repository, orchestrator, and agent execution suites. Each call restores an independent engine with no application tables or data. Every Chat migration and repository.bootstrap() still runs on that new engine, so these fixtures need no freshSchema bypass. See the exact proposed empty-engine helper and fixture specification.
The proposed shared fixture layer also uses that empty image in the central Bot and collaboration helpers. Collaboration callers still receive an empty schema and run their own bootstrap. Bot fixtures still bootstrap Chat on every call; migrate: false leaves Bot migrations for the caller. Application data, sequences, seeding, and teardown remain independent, and native PostgreSQL fixtures keep their existing behavior.
Linux benchmark images also need jq and the system Python cryptography package for existing workflow and host-script unit tests. Missing prerequisites are environment failures and must be resolved before comparing complete test results.