navmesh_explorer
Walks to random reachable points on the baked NavMesh, biased toward the least-visited 4 m cells, so side rooms and corners get visited. That's where level bugs hide.
AI-assisted game QA for Unity. A C# package records what happens during a playtest. A seeded bot plays the build while detectors watch for bugs. A Python tool turns thousands of log lines into a short list of unique, ranked bugs, with reports that cite their evidence.
A one-hour playtest of a Unity game can print thousands of warnings and errors. Most are repeats of a few real bugs, and the useful ones are buried. A QC tester then rewrites each one by hand into a ticket: what happened, where, how to reproduce it, how bad it is.
Studios automate parts of this with autoplay bots, crash grouping and, more and more, language models. But a model that invents reproduction steps or priorities makes things worse. The hard part is grouping well and staying honest about the evidence. That's what QA Lab is built around.
The Unity side and the Python side share only the JSON Schemas in schemas/. Both test against the same examples, so either side can change inside without breaking the other.
Add -qalab to a build, or tick auto-start in the editor. Log callbacks from any thread only enqueue; the main thread writes every 0.5 s, and a run always ends with run_end and no gaps in seq.
run.jsonBuild and commit, scenes, seed, clean end or crash, exit code.events.jsonlLogs with stack traces, bot actions, detector findings, screenshots, a metrics sample every second, markers.shots/Screenshots, long side ≤ 1280 px.results.xmlJUnit results for Jenkins and TeamCity.labels.jsonSeeded-bug and screenshot ground truth, benchmark mode only. Triage never opens it; a test enforces that.Every choice comes from a seeded RNG. Physics and timing still vary, so the action log is the repro record: each action becomes a step to reproduce.
Walks to random reachable points on the baked NavMesh, biased toward the least-visited 4 m cells, so side rooms and corners get visited. That's where level bugs hide.
Clicks random menu controls to shake out broken buttons, dead ends and placeholder UI.
A game adapter drives your game's own commands. A template ships as a package sample, including a turn-based policy.
A blocker or critical finding makes the player exit with code 1, so CI fails. A detector that throws is switched off for the run instead of taking the game down.
qalab vision analyze looks at the screenshots. Cheap pixel checks go first; a vision-language model is called only where they're blind: near UI actions and events, and on a sparse sample of other frames.
Findings join triage as visual bugs, with the clearest screenshot attached. The eval compares heuristics, the VLM, a logistic-regression baseline and the hybrid, per label.
Every line is checked against the JSON Schemas. Bad lines are reported, not silently skipped.
Ids and numbers come out of messages, stack frames are parsed, and events are grouped by a line-free signature. Four clustering variants: exact, frame_tfidf, frame_embed, tfidf_only.
score = weight × (1 + log2 count) × (1 + 0.5 × (runs − 1)) × crash. Priority P1–P4 follows from the score. The model never sets it.
An LLM drafts each report with context: nearby bot actions, logs and design-doc excerpts (RAG). JSON-schema output at temperature 0, pydantic validation, 2 retries, then a template fallback. Unknown ids are dropped and flagged.
report.html (offline, filterable), bugs.json, report.md and bugs_jira.csv. Exit code 3 when a P1 is found, so CI can fail on it.
A sandbox project has 16 seeded bugs with known causes, so grouping and detection can be scored against the truth instead of eyeballed.
scripts\benchmark.ps1Records one bot playtest per seed (20 by default), with labels.qalab eval triagePairwise precision, recall and F1 per clustering variant; report fallback rate, grounding, repro steps, retrieval hit@3, latency, tokens.qalab eval visionPer-label precision, recall and F1, false positives per 100 frames, VLM calls, latency and cost.Not measured yet. The evaluations are built; the benchmark needs the sandbox running in Unity. Numbers will appear here with the command that reproduces each one.
Windows PowerShell with Python 3.12+. The sample triage runs with a deterministic stand-in model, so you can see a full report before you set anything up.
python3 -m venv .venv && . .venv/bin/activate && pip install -e "./python[dev]"Linux / macOS: then the same commands with /
> git clone https://github.com/sorinnha/miaw-qa-lab > cd miaw-qa-lab > scripts\setup.ps1 # venv, install, .env, validate sample, run tests > .\.venv\Scripts\Activate.ps1
> qalab triage run samples\sample_run ` --provider fake ` --docs docs\sandbox_design.md ` --out out\sample_report > start out\sample_report\report.html
# once: open unity\QALabSandbox in Unity Hub, then Tools → QA Lab → Rebuild Sandbox Scenes > scripts\run_pipeline.ps1 -Seed 42 -Duration 120 -Open # build if needed → bot playtest + menu crawl → vision → triage → report.html
ui_overflow has no heuristic.--provider none writes template reports and retrieval falls back to TF-IDF. A model (Gemini hosted, or a local model through Ollama) improves the written summaries; grouping and priority never depend on it.none, everything stays local. API keys are read from the environment or .env and are never logged, cached or written to reports. Prompts are not logged.com.miawworks.qalab). The Python tools need Python 3.12+.results.xml for Jenkins and TeamCity. qalab triage exits with code 3 when it finds a P1. An example Jenkinsfile is included; it hasn't run on a real Jenkins server yet.inferred. If a draft fails, a template report is used.