Apache 2.0 Open Source Sub-10ms Semantic Cache Python 3.10+ Self-Healing ReAct Loop

Open-Source AI Terminal Copilot
Engineered for Developers

aikord-cli is a subscription-free, local-first alternative to GitHub Copilot CLI. It pairs local Small Language Models (SLMs) with DuckDB vector caching, dynamic cloud routing, and self-healing execution to turn plain English into rock-solid terminal commands.

$ pip install -e . && aikord suggest "find large files"
⚡

Sub-10ms Vector Cache

DuckDB VSS with HNSW cosine indexing and local MiniLM embeddings bypasses LLMs entirely for repeated or semantically close queries.

🔀

Dynamic Complexity Router

Lightweight MoE heuristic routes simple commands to free local SLMs (Ollama) and complex multi-pipe pipelines to DeepSeek or Groq.

🛡️

3-Layer Safety Defense

Guards against destructive data loss (rm -rf, DROP TABLE, force git pushes) through schema contracts and keyword filters.

🔄

ReAct Self-Healing Loop

When a command fails, aikord captures stderr, diagnoses the root cause, and autonomously retries with corrected syntax.

Quickstart (5-Minute Setup)

Get up and running with aikord-cli on Linux, macOS, or Windows in minutes.

# 1. Clone repository
git clone https://github.com/Apoorv2402/aikord-cli.git
cd aikord-cli

# 2. Create virtual environment & activate
python3 -m venv .venv
source .venv/bin/activate

# 3. Install in development mode
pip install -e .

# 4. Run your first query
aikord suggest "show memory usage in human readable format"
Pro Tip: Testing Suite Included
Install with pip install -e ".[dev]" to run pytest unit tests and evaluate your active LLM provider against 50 ground-truth examples.

Local SLM Setup (Ollama)

For 100% offline, privacy-first execution with zero monthly subscriptions, aikord pairs with Ollama.

  1. Download and launch Ollama on your system.
  2. Pull the recommended lightweight coding model:
    ollama pull qwen2.5-coder:7b
  3. By default, aikord-cli connects to http://localhost:11434/v1. To test local inference explicitly:
    aikord suggest "list all docker containers" --provider local

Cloud Provider Setup (DeepSeek / Groq / OpenAI)

When query complexity exceeds the local routing threshold (or when forced via --provider cloud), aikord escalates to cloud models using the standard OpenAI wire format.

Provider Model Default Setup Command Latency Profile
DeepSeek (Default) deepseek-chat aikord config set cloud_api_key "sk-..." ~1200ms, lowest cost ($0.27/M)
Groq (Ultra-Fast) llama-3.3-70b-versatile aikord config set groq_api_key "gsk-..." ~300ms TTFT, real-time speed
OpenAI gpt-4o-mini aikord config set cloud_endpoint "https://api.openai.com/v1" ~1400ms, high reliability

Core Architecture Deep-Dive

1. Terminal RAG & Context Injection

Traditional copilots fail because they guess the user's shell environment. aikord-cli implements zero-overhead context scraping prior to inference via src/core/context.py:

  • OS & Version: Normalizes linux, darwin, or windows and extracts distro tags (e.g. Ubuntu 22.04).
  • Active Shell Dialect: Probes $SHELL, $PSModulePath, and $COMSPEC to distinguish bash, zsh, fish, powershell, and cmd.
  • Working Directory & Git: Scrapes current path and runs non-blocking git status --porcelain for unstaged files and active branch.

2. DuckDB VSS Vector Cache

Unlike key-value caches that require identical query strings, aikord uses embedded vector caching with DuckDB's vss extension:

-- DuckDB HNSW Cosine Index Schema
CREATE TABLE query_cache (
    id          VARCHAR PRIMARY KEY,
    query_text  TEXT NOT NULL,
    embedding   FLOAT[384] NOT NULL,
    action_json JSON NOT NULL,
    hit_count   INTEGER DEFAULT 0
);

CREATE INDEX cache_hnsw_idx ON query_cache 
USING HNSW (embedding) WITH (metric = 'cosine');

Queries achieving a cosine similarity ≥ 0.92 return verified cached commands in 3ms - 8ms without network requests.

3. Dynamic Complexity Router

The router evaluates query complexity using four weighted heuristic factors:

Complexity Scoring Formula
Score = (TokenLength × 0.30) + (PipeCount × 0.25) + (DestructiveOp × 0.30) + (GitRepo × 0.15)

Scores < 0.40 route locally; scores ≥ 0.40 escalate to cloud LLMs.

4. 3-Layer Defense-in-Depth Safety

To eliminate dangerous accidental commands, aikord utilizes three independent safety layers:

  1. JSON Schema Contract: The LLM is constrained to output is_destructive: bool and risk_level: "low" | "medium" | "high" | "critical".
  2. Pydantic Model Sanitization: Strips accidental backticks and markdown fences.
  3. Deterministic Keyword Filter: Hardcoded regex check in is_command_destructive() for rm, del, drop, kill, format, dd, git reset --hard.

5. ReAct Self-Healing Loop

When --auto-fix is active, failed commands trigger an automated diagnostic cycle:

  • Reason: Submits failed command, exit code, and stderr back to the LLM.
  • Act: Generates a corrected command conforming to ErrorRemediation.
  • Cycle Detection: Hashes SHA256(command + stderr[:100]) to prevent repetitive looping.
  • Observe: Retries execution up to max_retries (default: 3).

6. Telemetry & Observability

Written to ~/.config/aikord-cli/telemetry.jsonl. Records wall-clock latency, streaming Time-to-First-Token (TTFT), input/output token counts, estimated USD costs, and privacy-preserving SHA256 hashed queries.

Interactive Terminal Playground

Test the suggestion pipeline, router, safety guards, and self-healing loop directly in your browser.

🎮 Live Pipeline Simulation Browser Sandbox

Select a preset scenario below or input custom natural language:

aikord terminal preview — monokai theme

CLI Command Reference

aikord suggest

Convert natural language into a verified shell command.

aikord suggest [OPTIONS] QUERY
Option Short Type Default Description
--provider -p string auto Override complexity router: local, cloud, ollama, deepseek, groq.
--auto-fix -f flag False Automatically invoke ReAct self-healing loop on command failure.
--no-cache flag False Bypass semantic DuckDB vector cache and query the LLM directly.
--run -r flag False Execute command immediately, skipping interactive confirmation menu.

aikord explain

Provides a detailed, flag-by-flag plain English explanation of any shell command.

aikord explain "tar -czf backup.tar.gz ./src"

aikord cache

Inspect or clear the semantic vector database.

# View cache statistics
aikord cache stats

# Clear all entries
aikord cache clear --yes

aikord config

View and manage settings saved to ~/.config/aikord-cli/config.json.

# Display current config with redacted API keys
aikord config show

# Set a configuration parameter
aikord config set cloud_api_key "sk-ant-..."
aikord config set routing_threshold "0.35"

aikord plugin

List all loaded plugin extensions from user and system plugin directories.

aikord plugin list

aikord eval

Benchmark your active model against the 50 ground-truth examples.

aikord eval --provider cloud --dataset evals/dataset.json --limit 50

Visual Config Generator

Customize your aikord setup visually and generate your config.json or .env file.

0.92
Cosine similarity cutoff for semantic cache hits (0.92 recommended).
0.40
Lower = more queries to cloud; Higher = keep more local.
3
30s
Auto-Fix Always Enabled
Run ReAct healing on failures without -f flag
~/.config/aikord-cli/config.json
Environment Export (.env)

Plugin Development Guide

Create plugins by subclassing PluginBase and dropping your .py file into ~/.config/aikord-cli/plugins/. No core modifications needed.

~/.config/aikord-cli/plugins/audit_logger.py
from src.plugins.base_plugin import PluginBase, PluginManifest
from src.agent.executor import ExecutionResult

class AuditLogPlugin(PluginBase):
    __manifest__ = PluginManifest(
        name="audit-logger",
        version="1.0.0",
        description="Records commands to company SIEM",
        author="DevOps Team",
        hooks=["post_execute"],
    )

    async def post_execute(self, result: ExecutionResult, action, context) -> None:
        with open("/var/log/aikord_audit.log", "a") as f:
            f.write(f"{context.username} | {action.command} | exit:{result.exit_code}\n")

Benchmarking & Evals

aikord-cli includes a formal benchmark harness (evals/run_evals.py) to evaluate model performance across 50 ground-truth scenarios:

Benchmark Metric Target Objective Significance
JSON Parse Rate ≥ 98.0% Ensures zero crashes from malformed model tokens.
Destructive Recall 100.0% Zero false negatives allowed on commands that delete or wipe data.
Median Latency (p50) < 2000ms Keeps the terminal experience crisp and responsive.
Tail Latency (p95) < 5000ms Maximum allowable latency for complex multi-pipe pipelines.

Troubleshooting & FAQ

Q: How do I resolve "Connection refused" when running local queries?

Ensure Ollama is running (ollama serve) and that you have pulled the model with ollama pull qwen2.5-coder:7b. Verify reachability with curl http://localhost:11434/api/tags.

Q: Does the vector cache consume large amounts of disk space?

No. DuckDB's compressed format with all-MiniLM-L6-v2 (384 dimensions) requires roughly 1.5 KB per cached command. 10,000 commands take less than 15 MB of disk space.

Q: What happens if a destructive command is misclassified by the LLM?

aikord's 3-Layer Defense-in-Depth intercepts it. Even if the LLM claims is_destructive: false, the deterministic keyword engine scans for matches against DESTRUCTIVE_PATTERNS and forces a safety confirmation prompt.

ESC