A self-improving agentic loop — an agent that iteratively executes, critiques, and refines its own output until it converges on a good answer.
LoopForge runs each task through a LangGraph state graph of specialized nodes: an Executor that uses tools to answer the task, a Critic that scores the output across multiple quality dimensions, a Refiner that improves it based on the critique, and a Meta node that stores strategy memory for future tasks. The loop continues until the output score crosses a convergence threshold, a hard iteration cap is reached, or the token budget runs out.
Input
│
▼
┌─────────────┐
│ Executor │ ← ReAct pattern (Thought → Action → Observation)
│ (Groq LLM) │ ← Tools: web search, calculator, yfinance, python repl
└──────┬──────┘
│
▼
┌─────────────┐
│ Critic │ ← Scores: factuality, completeness, clarity, task_alignment
│ (score/10) │ ← Weighted overall score
└──────┬──────┘
│
▼
┌─────────────┐ score ≥ 7.5 ──▶ ┌──────┐
│ Router │ ─────────────────────▶│ Meta │──▶ END
│ │ max iterations ──▶│ │
│ │ budget exceeded ─▶│ │
└──────┬──────┘ └──────┘
│ score < threshold
▼
┌─────────────┐
│ Refiner │ ← Improves output using specific critique reasoning
└──────┬──────┘
│
└────────────────────────────▶ Executor (next iteration)
- Self-improving loop — executor → critic → refiner cycle with configurable convergence threshold
- ReAct executor — Thought/Action/Observation reasoning pattern with real tool use
- Structured critic — 4-axis rubric scoring with few-shot examples and score anchors
- Loop memory — ChromaDB stores task strategies across runs
- Circuit breaker — stops early if scores decline for 2 consecutive iterations
- Token budget guard — stops gracefully (not a crash) if a task runs away on tokens
- LangFuse tracing — optional; per-iteration spans with scores, latency, and token counts (no-ops if unset)
| Layer | Technology |
|---|---|
| Orchestration | LangGraph |
| LLM (primary) | Groq — llama-3.1-8b-instant |
| LLM (fallback) | HuggingFace — Meta-Llama-3.1-8B-Instruct |
| Tools | Tavily (search), yfinance, calculator, sandboxed Python REPL |
| Vector memory | ChromaDB |
| Observability | LangFuse (optional) |
loopforge/
├── main.py # CLI entry point
├── requirements.txt
├── core/
│ ├── graph.py # LangGraph state graph
│ ├── state.py # GraphState TypedDict
│ ├── router.py # Conditional edge logic
│ └── nodes/
│ ├── executor.py # ReAct agent with tool use
│ ├── critic.py # Rubric scorer (4 axes)
│ ├── refiner.py # Critique-driven improver
│ └── meta.py # Strategy memory + status resolution
├── tools/
│ ├── search.py # Tavily web search
│ ├── calculator.py # AST-safe math evaluator
│ ├── python_repl.py # Sandboxed Python executor
│ └── yfinance_tool.py # Market data
├── memory/
│ └── chroma.py # ChromaDB client
└── observability/
└── langfuse_client.py # Optional trace + span management
- Python 3.11
- API keys (see below)
git clone https://github.com/Sahojit/Loop-Forge.git
cd Loop-Forge
cp .env.example .envEdit .env and fill in the keys you need (at minimum GROQ_API_KEY; TAVILY_API_KEY if you want the executor to use web search):
GROQ_API_KEY=gsk_... # console.groq.com/keys
HUGGINGFACE_API_KEY=hf_... # huggingface.co/settings/tokens (fallback LLM)
TAVILY_API_KEY=tvly-... # app.tavily.com (search tool)python3.11 -m venv .venv
source .venv/bin/activate # Windows: .venv\Scripts\activate
pip install -r requirements.txtpython3 main.py "explain quantum entanglement in simple terms"Or run it with no argument to be prompted for a task:
python3 main.pyOutput includes the final answer plus the run's status, iteration count, score history, and tokens used.
| Variable | Default | Description |
|---|---|---|
CONVERGENCE_THRESHOLD |
7.5 |
Minimum critic score to stop the loop |
MAX_ITERATIONS_DEFAULT |
5 |
Max executor→critic→refiner cycles per task |
TOKEN_BUDGET_PER_TASK |
8000 |
Max tokens before the run stops early (status: budget_exceeded) |
CHROMA_PERSIST_DIR |
./chroma_db |
Where strategy memory is stored |
LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY / LANGFUSE_HOST |
unset | Optional tracing — no-ops if not set |
Each run finishes in one of these states, visible in the CLI output as status:
| Status | Meaning |
|---|---|
converged |
Critic score reached CONVERGENCE_THRESHOLD |
max_iter_reached |
Hit max_iterations without converging, or scores declined for 2 consecutive iterations (circuit breaker) |
budget_exceeded |
Ran out of TOKEN_BUDGET_PER_TASK mid-run; returns the best partial output so far |
MIT