Skip to main content
Super-Goose — AI Agent Platform

The first self-evolving, voice-enabled, production-grade AI agent platform — 16 integrated tools, 5 specialist agents, overnight self-improvement

License16 Tools5 AgentsVoiceStage 5Docker
165K+Lines of Rust
~2,100Test Functions
182Commits Ahead
16Integrated Tools
5Specialist Agents
208Agent Tests
8Agent Phases
7Workspace Crates

What's New in v1.24.05

Extended thinking, prompt caching, 12 SOTA features, and a major dead-code cleanup.

Native Extended Thinking

Anthropic extended thinking support with --thinking CLI flag and configurable budget

Prompt Caching

80-90% input cost reduction with automatic cache_control injection for system and tool messages

12 SOTA Agent Features

Reflexion, Budget Enforcement, Output Guardrails, Code-Test-Fix, /model Hot-Switch, Adaptive Compaction, Session Search, Auto-Detection, Rate Limiting, Bookmarks, and more

4,454 Lines Dead Code Removed

Comprehensive audit pass removing unused code, fixing warnings, and aligning tests across 27 files

Key Features

How Super-Goose Works

From voice or text input to production-quality code — every step automated, reviewed, and traced.

Super-Goose End-to-End Pipeline
1

Plan

Your request enters the Planner, which breaks it into structured steps with dependencies

Goose Core, LangGraph
2

Assign

The ALMAS Team Coordinator assigns each step to the right specialist agent

5 agents: Architect, Developer, QA, Security, Deployer
3

Execute

Each agent executes in an isolated MicroVM sandbox with snapshot/restore

microsandbox, Arrakis
4

Scan

Every line of generated code passes through security scanning and AST-aware analysis

Semgrep, ast-grep, CrossHair
5

Review

The Coach/Player QA system reviews output in up to 3 adversarial cycles

Coach (precise), Player (creative)
6

Ship

Approved code gets a PR-Agent review and full observability traces

PR-Agent, Langfuse
7

Evolve

Overnight, the Self-Evolution Engine analyzes attempts and optimizes prompts

DSPy, Inspect AI, Mem0

The 16 Integrated Tools

Every tool is wired through the ResourceCoordinator with initialization locks, lazy loading, and atomic operations.

Super-Goose 16-Tool Ecosystem

Core Agent Layer

Goose CoreRust agent engine with MCP protocol, multi-provider LLM support
AiderAI pair programmer with 14 code editing strategies
LangGraphGraph-based workflows with checkpoint/resume and time-travel
OpenHandsSandboxed execution with browser automation
Pydantic-AIType-safe structured outputs with runtime validation
ConsciousVoice-first interface with emotion detection (8 emotions)

Self-Evolution Engine

DSPyBayesian prompt optimization (MIPROv2, GEPA)
Inspect AIUK AISI evaluation framework for scoring agent quality
Mem0Graph memory (Neo4j + Qdrant) with trajectory recall

Sandbox & Execution

microsandboxMicroVM isolation with <200ms boot, MCP-native
ArrakisVM snapshots for Language Agent Tree Search (LATS)

Governance & Observability

LangfuseDistributed tracing with token/cost/latency dashboards
SemgrepPolicy-as-code security scanning on every diff
PR-AgentAI-powered code review, test generation, changelogs

Code Quality & Verification

ast-grepAST-aware structural code search and refactoring
CrossHairFormal verification with Z3 symbolic execution

Why Stage 5?

Current AI coding agents operate at Stage 4 — a single model, static prompts, no memory. Super-Goose breaks through to Stage 5.

Stage 4 vs Stage 5 Comparison
DimensionStage 4 (Current)Stage 5 (Super-Goose)
ArchitectureSingle agent5 specialist roles + orchestrator
MemoryNone — starts fresh every timeGraph memory with trajectory recall
PromptsStatic foreverDSPy-compiled, improve nightly
QualityNo review — user sees all errorsCoach/Player adversarial review (3 cycles)
VoiceText onlyMoshi 7B speech-to-speech (<200ms)
SecurityNoneSemgrep + CrossHair on every diff
SandboxDocker containersMicroVM isolation (<200ms boot)
ObservabilityNoneLangfuse traces + OpenTelemetry spans

The ALMAS 5-Agent Team

Every task flows through five specialist agents in sequence, each with role-based access control (RBAC).

ALMAS 5 Specialist Agents
📐

Architect

Creates design docs, PLAN.md, architecture decisions

Read all, write docs only
💻

Developer

Writes code, runs builds, creates implementations

Full code access
🧪

QA Engineer

Runs tests, checks coverage ≥ 80%, validates quality

Read all, write tests only
🛡️

Security

Runs cargo audit, Semgrep scans, vulnerability checks

Read all, scan tools only
🚀

Deployer

Builds release artifacts, manages deployment

Build/deploy access only

Self-Evolution: The Overnight Gym

Super-Goose gets better over time — automatically. Every night, the Overnight Gym runs a full optimization cycle.

Overnight Gym Self-Evolution Cycle
1

Load

Pulls 50 benchmark tasks from SWE-bench Verified

2

Execute

Runs each task in an isolated microsandbox MicroVM

3

Score

Inspect AI evaluates every attempt with model-graded metrics

4

Optimize

DSPy GEPA analyzes patterns across all attempts, compiles better prompts

5

Remember

Mem0 stores successful trajectories as entity-relationship graphs

6

Improve

Week-over-week improvement tracked for regression detection

The result: 90%+ token savings through progressive 3-layer context disclosure, and measurably better prompts every week.

Coach/Player Quality Gate

Every output goes through a dual-model adversarial review before you see it.

The Player

Temperature 0.7 — Creative

Executes the task with full tool access. Generates code, tests, and documentation. Optimized for breadth and creativity.

vs

The Coach

Temperature 0.3 — Precise

Reviews against quality standards: compilation, tests, security, coverage ≥ 90%. Optimized for accuracy and rigor.

If rejected, the Player self-improves with Coach feedback and retries — up to 3 cycles. This catches errors that single-pass agents miss.

Conscious: Voice-First Interface

Super-Goose speaks. The Conscious voice AI companion provides real-time conversation with emotion detection.

🎙️
Moshi 7B Engine

Native speech-to-speech, no transcription step

⚡
<200ms Latency

Real-time conversation

😊
Emotion Detection

Wav2Vec2 tracking 8 emotions at 85-90% accuracy

🎭
13 Personality Profiles

Each with 20+ behavioral sliders

🔀
Intent Routing

CHAT (conversational) vs ACTION (execute via Goose)

📡
Device Control

SSH, 3D printing, network scanning

Voice Stats

Latency<200ms
Emotions8 tracked
Accuracy85-90%
Profiles13
Sliders20+ each

Enterprise Capabilities

⚙️

Orchestrator

Dependency graphs, parallel execution, priority scheduling, auto-retry

💾

Checkpointing

LangGraph SQLite snapshots. Resume from any state. Time-travel debugging.

🔗

Supply Chain

cosign signing, Syft SBOM, Trivy scanning, OpenSSF Scorecard

🔄

CI/CD

GitHub Actions pipelines. Smart change detection. All platforms.

Phase 6: Memory System

Cross-session context retention with 4 memory tiers, real vector embeddings, and graph memory — 127 memory tests passing.

TierPurposeDecay
WorkingShort-term LRU cache for current conversation0.70
EpisodicSession history and conversation context0.90
SemanticLong-term facts and knowledge (vector search)0.99
ProceduralLearned patterns and procedures0.98

Real Embeddings

Candle sentence-transformer (all-MiniLM-L6-v2, 384-dim) with hash fallback

Mem0 Dual-Write

Local memory + Neo4j/Qdrant graph memory when available

Auto Consolidation

Working → Episodic → Semantic promotion by importance & access

/memory Command

Stats, clear, and save subcommands — 127 tests passing

Phase 6.3 & 6.4: HITL + Benchmarks

Interactive session control with breakpoints, and a built-in benchmark framework for tracking agent quality.

Human-in-the-Loop

Interactive Breakpoints

Tool-level and pattern-based breakpoints, periodic check-ins, plan approval gates. /pause, /resume, /breakpoint, /inspect slash commands. Feedback injection mid-execution.

&

Agent Benchmarks

goose-bench Framework

6 builtin benchmark tasks, 5 evaluator criteria (Contains, Matches, ToolCalled, FileExists, Custom). Regression detection with configurable thresholds and performance deltas.

Phase 7: Task Graph & Skill Registry

LangGraph-inspired DAG execution engine and composable skill system with dependency resolution.

Task Graph Engine

DAG with typed nodes — ToolCall, Prompt, Conditional, SubGraph, ParallelFork, Join. Cycle detection (DFS), topological sort (Kahn's), Mermaid export.

Graph Executor

Node state tracking with Ready, Running, Completed, Failed states. Retry support with configurable max attempts. Automatic dependency resolution.

Skill Registry

10 skill categories, typed parameters, search/discovery API. 4 builtins: code-review, test-generator, docs-generator, security-audit.

Skill Pipelines

Chain skills with input/output mapping and validation. Recursive dependency resolution with circular dependency detection.

38 unit tests — 18 graph + 20 skill registry — all passing.

Phase 8: Extended Thinking & Agentic Swarms

Chain-of-thought deliberation with 7 thinking patterns, and multi-agent swarm orchestration with 6 routing strategies.

Extended Thinking

7 Thinking Patterns
Linear — Decompose → Plan → Execute
TreeOfThought — Explore multiple branches, pick best
ReAct — Reason → Act → Observe → Repeat
Reflexive — Plan → Execute → Reflect → Revise
StepBack — Abstract the problem first, then solve
LeastToMost — Solve simpler sub-problems first
SelfDebate — Generate arguments for/against, decide

Configurable budget: max steps, duration, tokens, depth. Confidence tracking with auto-termination. 12 unit tests.

+

Agentic Swarms

6 Routing Strategies
RoundRobin — Cycle through agents equally
LeastBusy — Route to agent with fewest active tasks
SkillBased — Match task requirements to agent skills
PerformanceBased — Route to highest-performing agent
Hybrid — Weighted: 40% skill + 30% performance + 30% load
Random — Random selection for load distribution

10 agent roles, BatchProcessor, inter-agent messaging, EMA performance tracking. 16 unit tests.

28 unit tests — 12 extended thinking + 16 swarm — all passing.

Downloads

Pre-built binaries for all platforms — CLI, Desktop, and Docker.

PlatformCLIDesktop
Dockerghcr.io/ghenghis/super-goose:v1.24.05—
Windowsgoose-x86_64-pc-windows-msvc.zipGoose-win32-x64.zip
macOS ARMgoose-aarch64-apple-darwin.tar.bz2Goose.dmg
macOS Intelgoose-x86_64-apple-darwin.tar.bz2Goose-intel.dmg
Linux x86goose-x86_64-unknown-linux-gnu.tar.bz2.deb / .rpm
Linux ARMgoose-aarch64-unknown-linux-gnu.tar.bz2—
Releases PageDocker Package

Docker

Pull from GitHub Container Registry — includes both CLI and server binaries.

Pull Image

docker pull ghcr.io/ghenghis/super-goose:latest

Or a specific version:

docker pull ghcr.io/ghenghis/super-goose:v1.24.05

Run CLI

docker run --rm -it ghcr.io/ghenghis/super-goose:latest goose

Interactive mode with the Goose CLI agent.

Run Server

docker run -d -p 3284:3284 ghcr.io/ghenghis/super-goose:latest goosed

API server on port 3284 for remote clients.

Image Details

Base: Debian SlimArch: linux/amd64Binaries: goose + goosedRegistry: GHCR

Infrastructure

Backend services run on Docker Compose — one command to start everything.

Neo4j

Graph memory for Mem0

:7474

Qdrant

Vector search for Mem0

:6333

PostgreSQL

Langfuse database

:5432

ClickHouse

Langfuse analytics

:8123

Redis

Langfuse cache

:6379

MinIO

Langfuse blob storage

:9000

Langfuse

Observability dashboard

:3000

OTel Collector

OpenTelemetry trace pipeline

:4317