Skip to content

Troubleshooting

Indexed by symptom. Start with the quick diagnostics, then the section matching what you are seeing.

First three commands

coding --health                               # what the system thinks
docker compose -f docker/docker-compose.yml ps  # what is actually running
curl -s localhost:3034/health/state | jq .    # the raw state document

If those three disagree with each other, that disagreement is the diagnosis — something is reading a stale copy.

By symptom

What you see Section
Install failed, coding not found Installation issues
Sessions not being logged Session logging issues
Containers missing or restarting Docker issues
The knowledge graph is empty or stale Knowledge base issues
A tool call was blocked unexpectedly Constraint monitor issues
Everything is slow Performance issues

The rule that saves the most time

Check the reporter before the thing reported. A grey or stale health badge means the process writing that state has stopped — the services it describes may be entirely fine.

Last resort

A complete reset exists in the Deep tier. Try targeted fixes first; a reset discards state that is usually not the cause.

Diagnosing in the right order

Nearly every confusing failure here comes from checking things in the wrong order, because the layers nest. Docker underpins the services; the coordinator reports on them; the dashboard and status line read the coordinator. A failure at any level makes everything above it look broken.

So: Docker, then the coordinator, then state freshness, then the individual service. Only the last of those means what it appears to mean.

The categories

Installation — almost always Docker not running, a shell not reloaded, or submodules that did not come down with the clone.

Session logging — a per-project monitor that is absent, or one that is alive but wedged. Those need different fixes, and the status line distinguishes them: no badge means healthy, a badge means unhealthy, and a project missing from the list entirely means no monitor at all.

Docker — a container that never started is a different problem from one that is restarting. docker ps distinguishes them and the container logs explain the second.

Knowledge base — an empty or stale graph usually means no extraction pass has run recently, rather than a broken store. The pass is asynchronous and takes 10–20 minutes.

Constraints — a blocked call is usually correct. When it is genuinely a false positive, the supported response is an explicit override naming the constraint, not rewording until the pattern stops matching.

Performance — check whether something is wedged before assuming load. A stalled process consumes a slot without consuming CPU, and looks like everything simply being slow.

Resetting everything

The full procedure is in the Deep tier. It is the last resort rather than a first move, because it discards caches, indexes and health state that are usually not the cause — and a reset that resolves a problem without identifying it is a reset you will perform again.

Common issues and solutions.

Quick Diagnostics

# System health check
./scripts/test-coding.sh

# LSL validation
node scripts/validate-lsl-config.js

# Docker services status
docker compose -f docker/docker-compose.yml ps

Installation Issues

Commands Not Found

# Reload shell configuration
source ~/.bashrc  # or ~/.zshrc

# Verify PATH
echo $PATH | grep coding

# If missing, reinstall
./install.sh --update-shell-config

Permission Errors

# Make scripts executable
chmod +x install.sh
chmod +x bin/*

# Run installer
./install.sh

MCP Servers Not Loading

# Check MCP configuration
cat ~/.config/Claude/claude_desktop_config.json

# Reinstall MCP config
bash -c 'source ./install.sh && setup_mcp_config'   # regenerate only the MCP config

# Check server logs
ls ~/.claude/logs/mcp*.log

Agent Drops Back to the Shell on Startup

If coding prints [tmux-wrapper] Creating tmux session: … and then returns straight to the shell prompt (no interactive agent), the agent process exited during startup before the session became interactive — usually a transient timing race (e.g. an MCP pre-flight gate such as VKB on port 8080 not yet ready). The wrapper now surfaces why instead of failing silently: it prints a ⚠️ Session … exited diagnostic with the tail of the captured launch output, and writes the full log to .logs/launch/<session>.log.

# Inspect the most recent failed launch
ls -t .logs/launch/*.log | head -1 | xargs tail -40

# Re-running usually clears a startup timing race
coding --claude

LSL Issues

LSL Files Not Generated

# Check if monitor is running
ps aux | grep enhanced-transcript-monitor

# Check ETM heartbeat via coordinator (Phase 33+: .health/*.json files are
# no longer written; coordinator's lsl slice is the source of truth)
curl -fs http://localhost:3034/health/state \
  | jq '.lsl | to_entries | map(select(.key | endswith(":coding")))'

# Restart monitor
coding --restart-monitor

Classification Not Working

# Check classification logs
ls -la .coding/history/logs/classification/

# Verify configuration
cat config/live-logging-config.json | jq '.embedding_classifier'

Recovery from Transcripts

# Batch recover LSL files
PROJECT_PATH=/path/to/project CODING_REPO=/path/to/coding \
  node scripts/batch-lsl-processor.js from-transcripts ~/.claude/projects/-path-to-project

# Recover specific date range
PROJECT_PATH=/path/to/project CODING_REPO=/path/to/coding \
  node scripts/batch-lsl-processor.js retroactive 2024-12-01 2024-12-03

Docker Issues

Container Won't Start

# First try: clean start (kills orphaned processes hogging ports)
coding --force

# Check logs if --force doesn't help
docker compose -f docker/docker-compose.yml logs coding-services

# Force rebuild
docker compose -f docker/docker-compose.yml build --no-cache
docker compose -f docker/docker-compose.yml up -d

Port Conflicts

# Recommended: force-clean all coding processes and restart
coding --force

# This kills supervisors, health monitors, all processes on coding ports,
# stops Docker containers, then proceeds with a clean startup.
# Combine with agent flags: coding --force --claude

# Manual investigation if --force doesn't resolve it
lsof -i :3848
kill $(lsof -ti :3848)

# Or change ports in .env.ports

Connection Refused

# Check container status
docker compose -f docker/docker-compose.yml ps

# Test health endpoint
curl -v http://localhost:3848/health

Volume Permission Issues

# Fix directory permissions
mkdir -p .data/knowledge-graph
chmod -R 755 .data/

# Restart
docker compose -f docker/docker-compose.yml restart

Knowledge Base Issues

VKB Won't Start

# Check if installed
ls -la integrations/memory-visualizer/

# If missing, initialize submodule
git submodule update --init --recursive integrations/memory-visualizer
cd integrations/memory-visualizer
npm install && npm run build

# Test viewer
vkb --debug

Missing Knowledge Export

# Check knowledge-export files
ls -la .data/knowledge-export/*.json

# If missing, graph database will create on next run

Constraint Monitor Issues

Dashboard Not Loading

# Check if services are running
lsof -i :3030
lsof -i :3031

# Start dashboard
cd integrations/constraint-monitor
PORT=3030 npm run dashboard

Hooks Not Firing

# Check hook configuration
cat ~/.claude/settings.json | jq '.hooks'

# Verify hook script exists
ls -la integrations/constraint-monitor/src/hooks/pre-tool-hook-wrapper.js

Performance Issues

High Memory Usage

# Increase Node.js memory
export NODE_OPTIONS="--max_old_space_size=4096"

# Restart with higher limit
NODE_OPTIONS="--max_old_space_size=4096" coding --claude

Slow Processing

# Process with lower priority
nice -n 19 node scripts/enhanced-transcript-monitor.js &

# Reduce monitoring frequency in config

Complete Reset

If installation is corrupted:

# Uninstall
./uninstall.sh

# Remove all data (WARNING: loses knowledge base)
rm -rf ~/.coding-tools/
rm -rf integrations/memory-visualizer/node_modules
rm -rf .data/knowledge-export/*.json

# Reinstall
./install.sh
source ~/.bashrc

Getting Help

Diagnostic Information

# Collect diagnostics
echo "=== System Info ===" > diagnostics.txt
uname -a >> diagnostics.txt
node --version >> diagnostics.txt
echo "PWD: $(pwd)" >> diagnostics.txt

echo -e "\n=== Environment ===" >> diagnostics.txt
env | grep -E "(USER|CODING|LSL)" >> diagnostics.txt

echo -e "\n=== Configuration ===" >> diagnostics.txt
cat .coding/var/validation-report.json 2>/dev/null >> diagnostics.txt

echo "Diagnostics collected in diagnostics.txt"

Support Resources