Verify & Repair¶
Checking that the installation actually works, and fixing it when it does not.
The check¶
The status line's [🏥●] badge, and the dashboard at localhost:3032 — it should read Healthy. For a pass over every subsystem:
./scripts/test-coding.sh # checks only, changes nothing
./scripts/test-coding.sh --interactive # offers each repair it finds
Diagnose in order¶
The failures nest, so checking out of order misdiagnoses them:
- Is Docker running? (learning tiers) Most failures are this.
- Is the coordinator reachable?
curl -s localhost:3034/health/state | jq .— a grey badge means nothing else on the dashboard can be trusted. - Is that state fresh? Older than ~3 minutes means the writer stopped, not the services.
- Only then, look at individual services.
The usual suspects¶
| Symptom | Usually |
|---|---|
| Everything red | Docker not running |
coding not found | Shell not reloaded since install |
| Missing tools in the agent | A container that never started |
| A project absent from the status line | No session monitor for it |
Full reset¶
The Deep tier has a complete-reset procedure. Try the targeted repairs first — a reset discards local state that is usually fine.
Verifying¶
coding-features status # which features are on — check THIS first
./scripts/test-coding.sh --interactive # a guided pass with repairs offered
curl -s localhost:3034/health/state | jq . # the raw truth
The three differ in what they can tell you. The first says what is supposed to be running (a service that is off by tier is not broken), the second walks through each subsystem and offers fixes, and the third is the document the dashboard renders — useful precisely when it and the status line disagree.
Why the order matters¶
Health is layered, and so are its failures. Checking a service before checking the thing that reports on it produces confident wrong answers:
- Docker down makes everything red, including things that are fine.
- The coordinator unreachable makes the dashboard grey — that is the reader having nothing to read, not the services being down.
- Stale state (older than roughly three minutes) means the process writing it stopped. The services it describes may be perfectly healthy; you are looking at a snapshot.
- A single unhealthy service is the only one of these that means what it appears to mean.
Common fixes¶
coding: command not found — the shell has not been reloaded since the install. source ~/.zshrc or open a new terminal.
Tools missing inside the agent — a container did not start. Check what is actually running before suspecting configuration; a service absent from docker ps never started, which is a Docker problem rather than a coding one.
A project missing from the status line — that project has no session monitor. This is different from having an unhealthy one, and the fix is different too.
Everything looks fine but nothing is happening — look for a wedged process. One that has stalled does not die: it answers ps, holds its ports, and stops doing work. Sample it rather than checking whether it exists.
Resetting¶
There is a complete-reset procedure in the Deep tier. Reach for it last: it discards local state — caches, indexes, health files — that is usually not the problem, and a reset that "fixes" an issue without identifying it tends to be a reset you perform repeatedly.
Test your installation and fix common issues.
Quick Health Check¶
Open the dashboard at localhost:3032; Run Verification re-checks everything on demand:

A healthy system reads Healthy with no critical issues. First confirm what is supposed to be running — a feature switched off by your tier is not a failure:
If you see red indicators, follow the troubleshooting steps below.
Automated Test Suite¶
The test script verifies all components:
# Check-only mode (default) — changes nothing, suggests each repair it would make
./scripts/test-coding.sh
# Interactive repair mode
./scripts/test-coding.sh --interactive
# Verbose output
./scripts/test-coding.sh --verbose
Test Categories¶
| Test | What It Checks |
|---|---|
| Prerequisites | Node.js, Git, jq, Docker versions |
| Commands | coding, coding-features, semantic in PATH |
| Configuration | .env, hooks, MCP config |
| Services | Docker containers or native processes |
| Connectivity | Health endpoints, port availability |
| Knowledge Base | Graph database accessibility |
Detailed Verification¶
1. Check Prerequisites¶
# Node.js 22 LTS or newer
node --version
# Git
git --version
# jq
jq --version
# Docker
docker --version
docker info
2. Verify Commands Available¶
# Should show help
coding --help
coding-features status
# Check locations
which coding
which coding-features
If commands not found, check your PATH:
3. Check Docker Services¶
# List running containers
docker compose -f ~/Agentic/coding/docker/docker-compose.yml ps
# Expected output shows these services running:
# - coding-services (main MCP servers; graphify runs file-based inside this container)
# - qdrant (vector database)
# - redis (cache / queues)
4. Test Health Endpoints¶
# Semantic Analysis MCP
curl -s http://localhost:3848/health | jq .
# Constraint Monitor MCP
curl -s http://localhost:3031/health | jq .
# obs-api (knowledge store, retrieval)
curl -s http://localhost:12436/health | jq .
# Health coordinator
curl -s http://localhost:3034/health | jq .
# LLM proxy
curl -s http://localhost:12435/health | jq .
# Health Dashboard
curl -s http://localhost:3032/health | jq .
5. Verify LSL Monitor¶
# Check ETM heartbeats via the coordinator (Phase 33+: .health/*.json files
# stopped being written; coordinator's lsl slice is now the source of truth)
curl -fs http://localhost:3034/health/state \
| jq '.lsl | to_entries | map(select(.key | endswith(":coding"))) | .[0].value | {status, lastBeat, projectName}'
# Check if monitor process is running
ps aux | grep -v grep | grep "enhanced-transcript-monitor"
6. Test Knowledge Base¶
# The live store answers (obs-api owns it)
curl -s 'http://localhost:12436/api/v1/entities?limit=1' | jq '.success'
# Where each project's knowledge is persisted
curl -s http://localhost:12436/api/kb/layout | jq
Common Issues and Fixes¶
Services Not Starting¶
Symptoms: Health check shows services as red/unavailable
Fix:
# Restart Docker services
docker compose -f docker/docker-compose.yml restart
# If that fails, rebuild
docker compose -f docker/docker-compose.yml down
docker compose -f docker/docker-compose.yml up -d --build
Port Conflicts¶
Symptoms: "Address already in use" errors
Fix:
# Find what's using the port (obs-api shown)
lsof -i :12436
# Kill the process or change port in .env.ports
nano .env.ports
# Restart services after port change
docker compose -f docker/docker-compose.yml down
docker compose -f docker/docker-compose.yml up -d
MCP Connection Errors¶
Symptoms: Claude can't connect to MCP servers
Fix:
# Verify MCP config
cat ~/.claude/settings.json | jq '.mcpServers'
# Reinstall hooks and config
./install.sh --mcp-only
LSL Not Recording¶
Symptoms: No files in .coding/history/
Fix:
# Check ETM heartbeat via coordinator (Phase 33+)
curl -fs http://localhost:3034/health/state \
| jq '.lsl | to_entries | map(select(.key | endswith(":coding")))'
# Restart monitor
coding --restart-monitor
# If still failing, check logs
tail -100 .logs/transcript-monitor-test.log
Knowledge Base Corrupted¶
Symptoms: obs-api will not start or crash-loops on its store, the viewer shows errors, the knowledge-base workflow fails.
The live store (LevelDB) is a machine-local cache: every project's knowledge is persisted in its repo's .coding/kb/ (see Per-Repo Tenancy), and a fresh store hydrates from all of them. So the fix is to rebuild the cache, never to delete .coding/kb/.
Fix:
# 1. Stop obs-api — the only process that owns the store
launchctl bootout gui/$(id -u)/com.coding.obs-api # macOS
systemctl --user stop obs-api # Linux / WSL
# 2. Move the store aside (keep it until the rebuild is verified)
DATA_HOME=$(~/Agentic/coding/bin/coding-data-home)
mv "$DATA_HOME/var/knowledge-graph/leveldb" "$DATA_HOME/var/knowledge-graph/leveldb.broken"
# 3. Start obs-api again — it hydrates from every .coding/kb/ it can see
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.coding.obs-api.plist # macOS
systemctl --user start obs-api # Linux / WSL
# 4. Check, then re-embed what is missing from Qdrant (idempotent)
curl -s 'http://localhost:12436/api/v1/entities?limit=1' | jq '.success'
node ~/Agentic/coding/dist/embedding/backfill.js
Repair Commands¶
Full Repair¶
This will: 1. Check all prerequisites 2. Reinstall missing components 3. Rebuild Docker containers 4. Reset configuration to defaults 5. Restart all services
Selective Repair¶
# Repair only MCP configuration
./install.sh --mcp-only
# Repair only hooks
./install.sh --hooks-only
# Rebuild Docker containers
cd docker && docker-compose build --no-cache && docker-compose up -d
Reset to Clean State¶
Data Loss
This removes all local data including knowledge base and session logs.
# Stop everything
docker compose -f docker/docker-compose.yml down 2>/dev/null
pkill -f "coding"
# Remove state files
rm -rf .data .health .logs .cache
rm -f .transition-in-progress
# Reinstall
./install.sh
Diagnostic Information¶
When reporting issues, include this diagnostic output:
# Generate diagnostic report
coding --diagnostics > diagnostics.txt
# Or manually collect:
echo "=== System ===" >> diagnostics.txt
uname -a >> diagnostics.txt
echo "=== Node ===" >> diagnostics.txt
node --version >> diagnostics.txt
echo "=== Docker ===" >> diagnostics.txt
docker --version >> diagnostics.txt
docker compose version >> diagnostics.txt
echo "=== Services ===" >> diagnostics.txt
docker compose -f docker/docker-compose.yml ps 2>&1 >> diagnostics.txt
echo "=== Health ===" >> diagnostics.txt
coding-features status >> diagnostics.txt 2>&1
curl -s http://localhost:3034/health/state >> diagnostics.txt 2>&1
Related Documentation¶
- Installation Guide - Fresh installation
- Configuration - API keys and settings
- Troubleshooting Reference - Extended troubleshooting