The first self-evolving, voice-enabled, production-grade AI agent platform — 16 integrated tools, 5 specialist agents, overnight self-improvement
Extended thinking, prompt caching, 12 SOTA features, and a major dead-code cleanup.
Anthropic extended thinking support with --thinking CLI flag and configurable budget
80-90% input cost reduction with automatic cache_control injection for system and tool messages
Reflexion, Budget Enforcement, Output Guardrails, Code-Test-Fix, /model Hot-Switch, Adaptive Compaction, Session Search, Auto-Detection, Rate Limiting, Bookmarks, and more
Comprehensive audit pass removing unused code, fixing warnings, and aligning tests across 27 files
From voice or text input to production-quality code — every step automated, reviewed, and traced.
Your request enters the Planner, which breaks it into structured steps with dependencies
Goose Core, LangGraphThe ALMAS Team Coordinator assigns each step to the right specialist agent
5 agents: Architect, Developer, QA, Security, DeployerEach agent executes in an isolated MicroVM sandbox with snapshot/restore
microsandbox, ArrakisEvery line of generated code passes through security scanning and AST-aware analysis
Semgrep, ast-grep, CrossHairThe Coach/Player QA system reviews output in up to 3 adversarial cycles
Coach (precise), Player (creative)Approved code gets a PR-Agent review and full observability traces
PR-Agent, LangfuseOvernight, the Self-Evolution Engine analyzes attempts and optimizes prompts
DSPy, Inspect AI, Mem0Every tool is wired through the ResourceCoordinator with initialization locks, lazy loading, and atomic operations.
Current AI coding agents operate at Stage 4 — a single model, static prompts, no memory. Super-Goose breaks through to Stage 5.
Every task flows through five specialist agents in sequence, each with role-based access control (RBAC).
Creates design docs, PLAN.md, architecture decisions
Read all, write docs onlyWrites code, runs builds, creates implementations
Full code accessRuns tests, checks coverage ≥ 80%, validates quality
Read all, write tests onlyRuns cargo audit, Semgrep scans, vulnerability checks
Read all, scan tools onlyBuilds release artifacts, manages deployment
Build/deploy access onlySuper-Goose gets better over time — automatically. Every night, the Overnight Gym runs a full optimization cycle.
Pulls 50 benchmark tasks from SWE-bench Verified
Runs each task in an isolated microsandbox MicroVM
Inspect AI evaluates every attempt with model-graded metrics
DSPy GEPA analyzes patterns across all attempts, compiles better prompts
Mem0 stores successful trajectories as entity-relationship graphs
Week-over-week improvement tracked for regression detection
Every output goes through a dual-model adversarial review before you see it.
Executes the task with full tool access. Generates code, tests, and documentation. Optimized for breadth and creativity.
Reviews against quality standards: compilation, tests, security, coverage ≥ 90%. Optimized for accuracy and rigor.
If rejected, the Player self-improves with Coach feedback and retries — up to 3 cycles. This catches errors that single-pass agents miss.
Super-Goose speaks. The Conscious voice AI companion provides real-time conversation with emotion detection.
Native speech-to-speech, no transcription step
Real-time conversation
Wav2Vec2 tracking 8 emotions at 85-90% accuracy
Each with 20+ behavioral sliders
CHAT (conversational) vs ACTION (execute via Goose)
SSH, 3D printing, network scanning
Dependency graphs, parallel execution, priority scheduling, auto-retry
LangGraph SQLite snapshots. Resume from any state. Time-travel debugging.
cosign signing, Syft SBOM, Trivy scanning, OpenSSF Scorecard
GitHub Actions pipelines. Smart change detection. All platforms.
Cross-session context retention with 4 memory tiers, real vector embeddings, and graph memory — 127 memory tests passing.
| Tier | Purpose | Decay |
|---|---|---|
| Working | Short-term LRU cache for current conversation | 0.70 |
| Episodic | Session history and conversation context | 0.90 |
| Semantic | Long-term facts and knowledge (vector search) | 0.99 |
| Procedural | Learned patterns and procedures | 0.98 |
Candle sentence-transformer (all-MiniLM-L6-v2, 384-dim) with hash fallback
Local memory + Neo4j/Qdrant graph memory when available
Working → Episodic → Semantic promotion by importance & access
Stats, clear, and save subcommands — 127 tests passing
Interactive session control with breakpoints, and a built-in benchmark framework for tracking agent quality.
Tool-level and pattern-based breakpoints, periodic check-ins, plan approval gates. /pause, /resume, /breakpoint, /inspect slash commands. Feedback injection mid-execution.
6 builtin benchmark tasks, 5 evaluator criteria (Contains, Matches, ToolCalled, FileExists, Custom). Regression detection with configurable thresholds and performance deltas.
LangGraph-inspired DAG execution engine and composable skill system with dependency resolution.
DAG with typed nodes — ToolCall, Prompt, Conditional, SubGraph, ParallelFork, Join. Cycle detection (DFS), topological sort (Kahn's), Mermaid export.
Node state tracking with Ready, Running, Completed, Failed states. Retry support with configurable max attempts. Automatic dependency resolution.
10 skill categories, typed parameters, search/discovery API. 4 builtins: code-review, test-generator, docs-generator, security-audit.
Chain skills with input/output mapping and validation. Recursive dependency resolution with circular dependency detection.
Chain-of-thought deliberation with 7 thinking patterns, and multi-agent swarm orchestration with 6 routing strategies.
Configurable budget: max steps, duration, tokens, depth. Confidence tracking with auto-termination. 12 unit tests.
10 agent roles, BatchProcessor, inter-agent messaging, EMA performance tracking. 16 unit tests.
Pre-built binaries for all platforms — CLI, Desktop, and Docker.
Pull from GitHub Container Registry — includes both CLI and server binaries.
docker pull ghcr.io/ghenghis/super-goose:latestOr a specific version:
docker pull ghcr.io/ghenghis/super-goose:v1.24.05docker run --rm -it ghcr.io/ghenghis/super-goose:latest gooseInteractive mode with the Goose CLI agent.
docker run -d -p 3284:3284 ghcr.io/ghenghis/super-goose:latest goosedAPI server on port 3284 for remote clients.
Backend services run on Docker Compose — one command to start everything.
Graph memory for Mem0
:7474Vector search for Mem0
:6333Langfuse database
:5432Langfuse analytics
:8123Langfuse cache
:6379Langfuse blob storage
:9000Observability dashboard
:3000OpenTelemetry trace pipeline
:4317