Skip to content

Verify & Repair

Checking that the installation actually works, and fixing it when it does not.

The check

The status line's [🏥●] badge, and the dashboard at localhost:3032 — it should read Healthy. For a pass over every subsystem:

./scripts/test-coding.sh                 # checks only, changes nothing
./scripts/test-coding.sh --interactive   # offers each repair it finds

Diagnose in order

The failures nest, so checking out of order misdiagnoses them:

  1. Is Docker running? (learning tiers) Most failures are this.
  2. Is the coordinator reachable? curl -s localhost:3034/health/state | jq . — a grey badge means nothing else on the dashboard can be trusted.
  3. Is that state fresh? Older than ~3 minutes means the writer stopped, not the services.
  4. Only then, look at individual services.

The usual suspects

Symptom Usually
Everything red Docker not running
coding not found Shell not reloaded since install
Missing tools in the agent A container that never started
A project absent from the status line No session monitor for it

Full reset

The Deep tier has a complete-reset procedure. Try the targeted repairs first — a reset discards local state that is usually fine.

Verifying

coding-features status                   # which features are on — check THIS first
./scripts/test-coding.sh --interactive   # a guided pass with repairs offered
curl -s localhost:3034/health/state | jq .   # the raw truth

The three differ in what they can tell you. The first says what is supposed to be running (a service that is off by tier is not broken), the second walks through each subsystem and offers fixes, and the third is the document the dashboard renders — useful precisely when it and the status line disagree.

Why the order matters

Health is layered, and so are its failures. Checking a service before checking the thing that reports on it produces confident wrong answers:

  • Docker down makes everything red, including things that are fine.
  • The coordinator unreachable makes the dashboard grey — that is the reader having nothing to read, not the services being down.
  • Stale state (older than roughly three minutes) means the process writing it stopped. The services it describes may be perfectly healthy; you are looking at a snapshot.
  • A single unhealthy service is the only one of these that means what it appears to mean.

Common fixes

coding: command not found — the shell has not been reloaded since the install. source ~/.zshrc or open a new terminal.

Tools missing inside the agent — a container did not start. Check what is actually running before suspecting configuration; a service absent from docker ps never started, which is a Docker problem rather than a coding one.

A project missing from the status line — that project has no session monitor. This is different from having an unhealthy one, and the fix is different too.

Everything looks fine but nothing is happening — look for a wedged process. One that has stalled does not die: it answers ps, holds its ports, and stops doing work. Sample it rather than checking whether it exists.

Resetting

There is a complete-reset procedure in the Deep tier. Reach for it last: it discards local state — caches, indexes, health files — that is usually not the problem, and a reset that "fixes" an issue without identifying it tends to be a reset you perform repeatedly.

Test your installation and fix common issues.


Quick Health Check

Open the dashboard at localhost:3032; Run Verification re-checks everything on demand:

System Healthy

A healthy system reads Healthy with no critical issues. First confirm what is supposed to be running — a feature switched off by your tier is not a failure:

coding-features status

If you see red indicators, follow the troubleshooting steps below.


Automated Test Suite

The test script verifies all components:

# Check-only mode (default) — changes nothing, suggests each repair it would make
./scripts/test-coding.sh

# Interactive repair mode
./scripts/test-coding.sh --interactive

# Verbose output
./scripts/test-coding.sh --verbose

Test Categories

Test What It Checks
Prerequisites Node.js, Git, jq, Docker versions
Commands coding, coding-features, semantic in PATH
Configuration .env, hooks, MCP config
Services Docker containers or native processes
Connectivity Health endpoints, port availability
Knowledge Base Graph database accessibility

Detailed Verification

1. Check Prerequisites

# Node.js 22 LTS or newer
node --version

# Git
git --version

# jq
jq --version

# Docker
docker --version
docker info

2. Verify Commands Available

# Should show help
coding --help
coding-features status

# Check locations
which coding
which coding-features

If commands not found, check your PATH:

echo $PATH | grep -q "Agentic/coding/bin" && echo "OK" || echo "Missing from PATH"

3. Check Docker Services

# List running containers
docker compose -f ~/Agentic/coding/docker/docker-compose.yml ps

# Expected output shows these services running:
# - coding-services (main MCP servers; graphify runs file-based inside this container)
# - qdrant (vector database)
# - redis (cache / queues)

4. Test Health Endpoints

# Semantic Analysis MCP
curl -s http://localhost:3848/health | jq .

# Constraint Monitor MCP
curl -s http://localhost:3031/health | jq .

# obs-api (knowledge store, retrieval)
curl -s http://localhost:12436/health | jq .

# Health coordinator
curl -s http://localhost:3034/health | jq .

# LLM proxy
curl -s http://localhost:12435/health | jq .

# Health Dashboard
curl -s http://localhost:3032/health | jq .

5. Verify LSL Monitor

# Check ETM heartbeats via the coordinator (Phase 33+: .health/*.json files
# stopped being written; coordinator's lsl slice is now the source of truth)
curl -fs http://localhost:3034/health/state \
  | jq '.lsl | to_entries | map(select(.key | endswith(":coding"))) | .[0].value | {status, lastBeat, projectName}'

# Check if monitor process is running
ps aux | grep -v grep | grep "enhanced-transcript-monitor"

6. Test Knowledge Base

# The live store answers (obs-api owns it)
curl -s 'http://localhost:12436/api/v1/entities?limit=1' | jq '.success'

# Where each project's knowledge is persisted
curl -s http://localhost:12436/api/kb/layout | jq

Common Issues and Fixes

Services Not Starting

Symptoms: Health check shows services as red/unavailable

Fix:

# Restart Docker services
docker compose -f docker/docker-compose.yml restart

# If that fails, rebuild
docker compose -f docker/docker-compose.yml down
docker compose -f docker/docker-compose.yml up -d --build

Port Conflicts

Symptoms: "Address already in use" errors

Fix:

# Find what's using the port (obs-api shown)
lsof -i :12436

# Kill the process or change port in .env.ports
nano .env.ports

# Restart services after port change
docker compose -f docker/docker-compose.yml down
docker compose -f docker/docker-compose.yml up -d

MCP Connection Errors

Symptoms: Claude can't connect to MCP servers

Fix:

# Verify MCP config
cat ~/.claude/settings.json | jq '.mcpServers'

# Reinstall hooks and config
./install.sh --mcp-only

LSL Not Recording

Symptoms: No files in .coding/history/

Fix:

# Check ETM heartbeat via coordinator (Phase 33+)
curl -fs http://localhost:3034/health/state \
  | jq '.lsl | to_entries | map(select(.key | endswith(":coding")))'

# Restart monitor
coding --restart-monitor

# If still failing, check logs
tail -100 .logs/transcript-monitor-test.log

Knowledge Base Corrupted

Symptoms: obs-api will not start or crash-loops on its store, the viewer shows errors, the knowledge-base workflow fails.

The live store (LevelDB) is a machine-local cache: every project's knowledge is persisted in its repo's .coding/kb/ (see Per-Repo Tenancy), and a fresh store hydrates from all of them. So the fix is to rebuild the cache, never to delete .coding/kb/.

Fix:

# 1. Stop obs-api — the only process that owns the store
launchctl bootout gui/$(id -u)/com.coding.obs-api          # macOS
systemctl --user stop obs-api                              # Linux / WSL

# 2. Move the store aside (keep it until the rebuild is verified)
DATA_HOME=$(~/Agentic/coding/bin/coding-data-home)
mv "$DATA_HOME/var/knowledge-graph/leveldb" "$DATA_HOME/var/knowledge-graph/leveldb.broken"

# 3. Start obs-api again — it hydrates from every .coding/kb/ it can see
launchctl bootstrap gui/$(id -u) ~/Library/LaunchAgents/com.coding.obs-api.plist   # macOS
systemctl --user start obs-api                                                      # Linux / WSL

# 4. Check, then re-embed what is missing from Qdrant (idempotent)
curl -s 'http://localhost:12436/api/v1/entities?limit=1' | jq '.success'
node ~/Agentic/coding/dist/embedding/backfill.js


Repair Commands

Full Repair

./install.sh --repair

This will: 1. Check all prerequisites 2. Reinstall missing components 3. Rebuild Docker containers 4. Reset configuration to defaults 5. Restart all services

Selective Repair

# Repair only MCP configuration
./install.sh --mcp-only

# Repair only hooks
./install.sh --hooks-only

# Rebuild Docker containers
cd docker && docker-compose build --no-cache && docker-compose up -d

Reset to Clean State

Data Loss

This removes all local data including knowledge base and session logs.

# Stop everything
docker compose -f docker/docker-compose.yml down 2>/dev/null
pkill -f "coding"

# Remove state files
rm -rf .data .health .logs .cache
rm -f .transition-in-progress

# Reinstall
./install.sh

Diagnostic Information

When reporting issues, include this diagnostic output:

# Generate diagnostic report
coding --diagnostics > diagnostics.txt

# Or manually collect:
echo "=== System ===" >> diagnostics.txt
uname -a >> diagnostics.txt
echo "=== Node ===" >> diagnostics.txt
node --version >> diagnostics.txt
echo "=== Docker ===" >> diagnostics.txt
docker --version >> diagnostics.txt
docker compose version >> diagnostics.txt
echo "=== Services ===" >> diagnostics.txt
docker compose -f docker/docker-compose.yml ps 2>&1 >> diagnostics.txt
echo "=== Health ===" >> diagnostics.txt
coding-features status >> diagnostics.txt 2>&1
curl -s http://localhost:3034/health/state >> diagnostics.txt 2>&1