Miaw QA Lab.

AI-assisted game QA for Unity. A C# package records what happens during a playtest. A seeded bot plays the build while detectors watch for bugs. A Python tool turns thousands of log lines into a short list of unique, ranked bugs, with reports that cite their evidence.

  • v0.1.0 in development
  • MIT license
  • Unity 2022.3+
  • Python 3.12+
  • Windows-first · runs on Linux & macOS
  • offline mode, no model needed
// the problem

Thousands of lines.
A handful of bugs.

A one-hour playtest of a Unity game can print thousands of warnings and errors. Most are repeats of a few real bugs, and the useful ones are buried. A QC tester then rewrites each one by hand into a ticket: what happened, where, how to reproduce it, how bad it is.

Studios automate parts of this with autoplay bots, crash grouping and, more and more, language models. But a model that invents reproduction steps or priorities makes things worse. The hard part is grouping well and staying honest about the evidence. That's what QA Lab is built around.

// architecture

Two halves.
One contract.

The Unity side and the Python side share only the JSON Schemas in schemas/. Both test against the same examples, so either side can change inside without breaking the other.

unity · c#

com.miawworks.qalab

  • LogCapture · MetricsSampler
  • EventWriter (thread-safe)
  • BotRunner + adapters
  • DetectorHub · Screenshots
  • ProjectScanner (editor)
run folder

runs/<run_id>/

  • run.json
  • events.jsonl
  • shots/*.png
  • results.xml (JUnit)
  • labels.json (benchmark)
python · qalab

qalab CLI

  • validate (JSON Schemas)
  • normalise → cluster → rank
  • RAG: design docs + code
  • LLM draft → grounding checks
  • vision · eval
outputs

reports

  • report.html (offline)
  • bugs.json
  • report.md
  • bugs_jira.csv
// 1-1 · record

Every run
leaves a folder.

Add -qalab to a build, or tick auto-start in the editor. Log callbacks from any thread only enqueue; the main thread writes every 0.5 s, and a run always ends with run_end and no gaps in seq.

  • run.jsonBuild and commit, scenes, seed, clean end or crash, exit code.
  • events.jsonlLogs with stack traces, bot actions, detector findings, screenshots, a metrics sample every second, markers.
  • shots/Screenshots, long side ≤ 1280 px.
  • results.xmlJUnit results for Jenkins and TeamCity.
  • labels.jsonSeeded-bug and screenshot ground truth, benchmark mode only. Triage never opens it; a test enforces that.

// 1-2 · play & 1-3 · watch

A bot that explores,
not wanders.

Every choice comes from a seeded RNG. Physics and timing still vary, so the action log is the repro record: each action becomes a step to reproduce.

adapter

navmesh_explorer

Walks to random reachable points on the baked NavMesh, biased toward the least-visited 4 m cells, so side rooms and corners get visited. That's where level bugs hide.

adapter

ui_crawler

Clicks random menu controls to shake out broken buttons, dead ends and placeholder UI.

adapter

your game

A game adapter drives your game's own commands. A template ships as a package sample, including a turn-based policy.

Detectors, each rate-limited per cell and attached to a screenshot

  • fell_out_of_world · and respawns the player
  • stuck
  • tunneling
  • perf_spike
  • exception_burst

A blocker or critical finding makes the player exit with code 1, so CI fails. A detector that throws is switched off for the run instead of taking the game down.

// 1-4 · see

Glitches you
can't grep.

qalab vision analyze looks at the screenshots. Cheap pixel checks go first; a vision-language model is called only where they're blind: near UI actions and events, and on a sparse sample of other frames.

missing_texturemagenta pixels · heuristic
black_screennear-black frame · heuristic
placeholder_uiwhite boxes · heuristic
ui_overflowtext spilling out · vision model

Findings join triage as visual bugs, with the clearest screenshot attached. The eval compares heuristics, the VLM, a logistic-regression baseline and the hybrid, per label.

// 1-5 · triage

From log spam
to tickets.

  1. step.1

    Validate

    Every line is checked against the JSON Schemas. Bad lines are reported, not silently skipped.

  2. step.2

    Normalise & group

    Ids and numbers come out of messages, stack frames are parsed, and events are grouped by a line-free signature. Four clustering variants: exact, frame_tfidf, frame_embed, tfidf_only.

  3. step.3

    Rank

    score = weight × (1 + log2 count) × (1 + 0.5 × (runs − 1)) × crash. Priority P1–P4 follows from the score. The model never sets it.

  4. step.4

    Draft & check

    An LLM drafts each report with context: nearby bot actions, logs and design-doc excerpts (RAG). JSON-schema output at temperature 0, pydantic validation, 2 retries, then a template fallback. Unknown ids are dropped and flagged.

  5. step.5

    Ship

    report.html (offline, filterable), bugs.json, report.md and bugs_jira.csv. Exit code 3 when a P1 is found, so CI can fail on it.

Pick your model, or none

geminiHosted. For sandbox data on the free tier.
ollamaLocal. Nothing leaves the machine.
noneTemplate reports. No model at all.
fakeDeterministic stand-in the tests use.
// 1-6 · measure

Seeds are
ground truth.

A sandbox project has 16 seeded bugs with known causes, so grouping and detection can be scored against the truth instead of eyeballed.

  • scripts\benchmark.ps1Records one bot playtest per seed (20 by default), with labels.
  • qalab eval triagePairwise precision, recall and F1 per clustering variant; report fallback rate, grounding, repro steps, retrieval hit@3, latency, tokens.
  • qalab eval visionPer-label precision, recall and F1, false positives per 100 frames, VLM calls, latency and cost.

Not measured yet. The evaluations are built; the benchmark needs the sandbox running in Unity. Numbers will appear here with the command that reproduces each one.

// quick start

Offline first.
No API key.

Windows PowerShell with Python 3.12+. The sample triage runs with a deterministic stand-in model, so you can see a full report before you set anything up.

python3 -m venv .venv && . .venv/bin/activate && pip install -e "./python[dev]"Linux / macOS: then the same commands with /

1 · install
> git clone https://github.com/sorinnha/miaw-qa-lab
> cd miaw-qa-lab
> scripts\setup.ps1
# venv, install, .env, validate sample, run tests
> .\.venv\Scripts\Activate.ps1
2 · triage the sample, offline
> qalab triage run samples\sample_run `
    --provider fake `
    --docs docs\sandbox_design.md `
    --out out\sample_report
> start out\sample_report\report.html
3 · a bot playtest of the sandbox (Windows + Unity)
# once: open unity\QALabSandbox in Unity Hub, then Tools → QA Lab → Rebuild Sandbox Scenes
> scripts\run_pipeline.ps1 -Seed 42 -Duration 120 -Open
# build if needed → bot playtest + menu crawl → vision → triage → report.html
// limitations

What it
can't do (yet).

  • No measured results yet. The triage and vision evaluations are built but have only run on copies of the hand-made sample.
  • The Unity code hasn't run in Unity yet. The engine-free C# is compiled and tested on .NET in CI; the rest was compiled against stand-in Unity types. First real runs are next.
  • Vision is tuned on synthetic frames. Post-processing like bloom can shift the magenta, and ui_overflow has no heuristic.
  • The bot is random. One seed may not reach every bug; detection rates come from the 20-seed benchmark.
  • Grouping is signature-based. It can split a bug whose message varies in ways normalisation misses.
  • Scale. Events are held in memory per run. No database or streaming for millions of events.

// faq

Questions
studios ask.

Does QA Lab need an LLM or an API key?
No. --provider none writes template reports and retrieval falls back to TF-IDF. A model (Gemini hosted, or a local model through Ollama) improves the written summaries; grouping and priority never depend on it.
Does my game's data leave my machine?
Only if you choose a hosted model. With Ollama or none, everything stays local. API keys are read from the environment or .env and are never logged, cached or written to reports. Prompts are not logged.
Which Unity versions are supported?
The package targets Unity 2022.3 or newer (UPM package com.miawworks.qalab). The Python tools need Python 3.12+.
Can it fail our CI build?
Yes. The player exits with code 1 on a blocker or critical detector finding and writes JUnit results.xml for Jenkins and TeamCity. qalab triage exits with code 3 when it finds a P1. An example Jenkinsfile is included; it hasn't run on a real Jenkins server yet.
How do you stop the AI from making things up?
Priority is computed by a formula, not chosen by the model. Output is JSON-schema constrained at temperature 0 and validated. Every evidence, action and doc id is checked against the run; unknown ids are dropped and flagged, and a repro step with no matching bot action is marked inferred. If a draft fails, a template report is used.
Is it open source?
Yes, MIT licensed, on GitHub.
Can we pilot it on our game?
Yes, that's what we're looking for. Email miaw@miaw-works.dev; we'll write a game adapter with you and you can keep everything on your own machines.
// pilot it

Try QA Lab
on your game.

miaw@miaw-works.dev