When a user hits a bug in Maestro, the last thing we want is for it to disappear into a support queue. We wanted a system where a user clicks “Report Issue,” and a few minutes later, a well-written GitHub issue appears in our repo – triaged, labeled, and ready for an engineer to pick up. No human in the loop unless the AI isn’t confident.
Here’s how we built it.
The problem
Maestro is a desktop application. When something goes wrong, users see the bug but we don’t – there’s no server-side telemetry to correlate with. Historically, users would describe the issue in Slack or email, we’d ask for logs, they’d send a partial paste, we’d ask for the database, and three days later we’d have enough context to start investigating.
We needed a way to collect rich diagnostics at the moment the bug occurs and route them to the right place automatically.
The architecture
The system has three stages: collect, analyze, and act.
User clicks "Report Issue"
|
v
Maestro client bundles diagnostics
(logs, config, SQLite DB, manifest)
|
v
POST /api/v1/issues --> Cloud Run service
| |
v v
GCS (zip archive) Firestore (metadata)
|
v
Eventarc (GCS finalization event)
|
v
Cloud Run Job (maestro-triage)
|
v
Claude Code analyzes diagnostics
|
v
GitHub issue created on SnapdragonPartners/maestro
Stage 1: Collect
When a user clicks “Report Issue” in Maestro’s web UI, a modal asks them to describe what happened and optionally include diagnostics.

If they opt in, the client bundles:
- maestro.log – the application log
- maestro.db – the SQLite database (checkpointed with
PRAGMA wal_checkpoint(PASSIVE)for consistency) - config.json – application configuration
- manifest.json – structured metadata including the Maestro version, installation ID, file sizes, and the user’s description
Sensitive files like encrypted secrets and password verifiers are explicitly excluded. The bundle is zipped and uploaded to our issue service.
The upload is signed with an HMAC derived from the installation ID and a shared secret embedded in the Maestro binary. This isn’t a security boundary – it’s an anti-abuse gate that prevents casual bot submissions without requiring user accounts.
Stage 2: Analyze
The issue service is a Go application running on Cloud Run. It validates the HMAC signature, checks rate limits (10 submissions per installation per 24 hours), assigns a sequential ID like MAE-0004, uploads the zip to GCS, and writes the metadata to Firestore with status new.
When the zip lands in GCS, an Eventarc trigger fires. This hits an endpoint on our service that kicks off a Cloud Run Job – the triage agent.
The triage agent is a Go orchestrator that:
- Claims the issue atomically (Firestore compare-and-set from
newtoin_flight, preventing duplicate processing) - Downloads and extracts the diagnostics zip, with zip-slip protection against path traversal
- Clones our repo twice: once at the version tag the user was running, once at HEAD (to check if the bug is already fixed)
- Builds a prompt from a Go template, injecting the issue metadata and file paths
- Invokes Claude Code in headless mode with the prompt
Claude gets access to Bash, Read, Write, Grep, and Glob tools. It searches the logs for errors, queries the SQLite database, checks existing GitHub issues for duplicates using the gh CLI, and reads the relevant source code. It writes a structured verdict to /tmp/verdict.json classifying the issue as valid, duplicate, invalid, or needs_review with a confidence score.
Here’s a real verdict from our first end-to-end run. A user reported that they couldn’t copy logs from the demo mode UI. Claude found the root cause in under three minutes:
The poll fires every ~3 seconds, so a user can never hold a selection long enough to copy.
logsContent.textContent = data.logs– this full DOM replacement destroys any active browser text selection on every poll cycle.
Classification: valid, confidence: 0.85.
Stage 3: Act
If the verdict is valid and confidence is at or above 0.7, the orchestrator creates a GitHub issue on our repo. The issue includes the summary, detailed analysis, evidence, and triage metadata. Labels are proposed by Claude and filtered against the repo’s actual label set before being applied.
If confidence is below 0.7, the issue is flagged as needs_review for a human to look at. Duplicates are linked to the existing issue. Invalid reports are recorded but no action is taken.
Here’s a real example of an auto-created issue from this pipeline.
The whole flow – from the user clicking “Report Issue” to a labeled GitHub issue appearing in our repo – takes about three to four minutes.
The hard parts
Sandboxing the AI
Claude has Bash access so it can run grep, sqlite3, git log, and gh issue list during analysis. That same Bash access means a malicious issue description could theoretically try to convince Claude to modify code or push changes.
We mitigate this with three independent layers, none of which rely on prompt compliance:
-
Environment allowlist: Claude’s subprocess only receives
ANTHROPIC_API_KEY,PATH,HOME, locale, and proxy variables. GitHub tokens, SSH agent sockets, cloud credentials – all stripped. The orchestrator retains them for its own use after Claude exits. -
Read-only repositories: After cloning, the orchestrator removes write permissions from both repo directories recursively. Symlinks are skipped to prevent chmod from escaping the tree.
-
No mutation tools: Claude’s prompt explicitly restricts it to analysis. Its only write target is
/tmp/verdict.json.
Even if all three fail, the damage is bounded: the container is ephemeral, has no persistent state, and the orchestrator validates the verdict structure before acting on it.
Sanitization
Everything Claude produces that becomes public (the GitHub issue title, body, and labels) passes through deterministic sanitization. We strip:
- The user’s installation ID (replaced with
[INSTALLATION_ID]) - UUIDs, IP addresses, and email addresses
- Container-internal paths like
/tmp/workspace/that leak infrastructure details - Patterns that look like API keys, tokens, or passwords
This is regex-based and runs after Claude’s analysis, so it catches anything the model might inadvertently include. It’s defense in depth – the prompt also instructs Claude to redact secrets, but we don’t rely on that alone.
Idempotency
Every step is designed to be safely retriable:
- Issue claiming uses a Firestore transaction that atomically transitions
newtoin_flight. If two workers try to claim the same issue, exactly one succeeds. - GitHub issue creation searches for existing issues containing the MAE ID before creating a new one.
- Status updates are the last step, so a crashed job leaves the issue in
in_flight(detectable) rather than in an inconsistent state.
The version tag problem
Maestro’s version string during development is dev, which doesn’t correspond to any git tag. Rather than failing, the orchestrator falls back gracefully: it clones HEAD and symlinks the version directory to it. Claude still gets two “repos” to compare, they’re just identical – and it correctly notes “not fixed on HEAD” since the code is the same.
Infrastructure
The system runs entirely on Google Cloud:
- Cloud Run service (the issue API): Go binary in a distroless container, scales to zero
- Cloud Run Job (the triage agent): Node.js base image with Go binary, sqlite3, git, GitHub CLI, and Claude Code installed. 2 vCPU, 2GB RAM, 1-hour timeout
- Eventarc: Wires GCS object finalization to the triage trigger endpoint
- Firestore: Issue metadata and verdicts
- GCS: Diagnostic zip archives
- Secret Manager: API keys for Anthropic and GitHub
CI/CD is two GitHub Actions workflows with workload identity federation – no long-lived service account keys. The triage image only rebuilds when relevant code changes.
What we learned
Let the orchestrator own all mutations. Claude analyzes. The Go code decides what to do with the analysis. This makes the system easier to reason about, easier to test, and safer. If we want to add a Slack notification or update a dashboard, we add it to the orchestrator, not to Claude’s prompt.
Allowlists beat denylists for security. Our first pass stripped GITHUB_TOKEN and GH_TOKEN from Claude’s environment. A reviewer pointed out that SSH_AUTH_SOCK, AWS_SECRET_ACCESS_KEY, and countless other variables could also leak credentials. Flipping to an allowlist of ~20 known-safe variables was simpler and more robust.
Start with the gate closed. AUTO_CREATE_ISSUES defaults to false. We ran several triage cycles with it off, inspecting verdicts in Firestore, before flipping it to true. This let us build confidence in the system before it started taking public actions.
Phased rollout works. Phase 1 was analysis-only (write verdict to Firestore). Phase 2 added GitHub issue creation. Phase 3 will add richer duplicate detection and subsystem routing. Each phase is a small, reviewable PR that builds on the last.
What’s next
We’re planning to add:
- User notifications: Letting submitters know their issue was triaged and linking them to the GitHub issue
- Subsystem labeling: Automatically routing issues to the right team based on which code paths are involved
- Feedback loop: Tracking whether auto-created issues get closed as “won’t fix” or “not a bug” and using that signal to improve the prompt
- Richer duplicate detection: Using embedding similarity in addition to keyword search
The code is in our maestro-issues repository if you want to see the implementation details.