Problem
My home game-streaming setup (gaming PC as host, a mini PC as client, a TV as display) showed occasional micro-stutter. Diagnosing it meant reading large host JSON exports and client logs by hand, or pasting them into a chatbot and hoping the numbers were right.
Approach
- Local-first. A static web app with no backend. Files are processed in the browser and never leave the device.
- Three sources, one timeline. Parsers for host session JSON (Sunshine-based host), client logs (Moonlight-based clients), and Steam Remote Play logs. Clocks are aligned from file-name epochs plus a measured residual offset.
- Correct counters. Host counters are cumulative per connection and reset on reconnect, so the app sums per-connection deltas instead of reading raw totals.
- Deterministic statistics. Average, P50, P5, and P1 FPS with interpolated percentiles; share of time at target FPS; bitrate and encode latency P95.
- Rule-based diagnostics. A median FPS baseline, episode grouping, and rules that separate "host overloaded" from "nothing to send" from loading screens. Gameplay is detected automatically from rolling CPU, FPS, and bitrate features.
- Explainable scores. 1–10 scores for FPS, stability, latency, network, and image, each shown with the numbers it was built from.
- Automation. A Python agent (standard library only) collects new logs and serves the app on the home network, so a phone sees the same sessions.
How AI was used
I drafted the specification with ChatGPT, kept a handoff document of decisions and confirmed file formats, and built the app with Claude Code. One rule from the start: AI writes code, but never produces the numbers. Every statistic is computed in plain JavaScript. The diagnostic rules were checked against real session files before being kept.
Core of the statistics
// Linear-interpolated quantile, used for P1 / P5 / P50 / P95
const quantile = (a, p) => {
if (!a.length) return null;
const s = sorted(a), i = (s.length - 1) * p,
lo = Math.floor(i), hi = Math.ceil(i);
return s[lo] + (s[hi] - s[lo]) * (i - lo);
};
Outcome
One drag-and-drop now gives a benchmark for a chosen stretch of gameplay, a diagnosis of the whole session, and a short report ready to share. Sessions can be saved and compared over time. About 3,300 lines of JavaScript and Python, open source.
Why it is on a BI CV: messy sources, metric definitions, reconciliation of counters, and outputs a non-expert can trust. It is the same work as BI, in a different domain.