Skip to content

Latest commit

 

History

51 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

LoopForge

A self-improving agentic loop — an agent that iteratively executes, critiques, and refines its own output until it converges on a good answer.


Overview

LoopForge runs each task through a LangGraph state graph of specialized nodes: an Executor that uses tools to answer the task, a Critic that scores the output across multiple quality dimensions, a Refiner that improves it based on the critique, and a Meta node that stores strategy memory for future tasks. The loop continues until the output score crosses a convergence threshold, a hard iteration cap is reached, or the token budget runs out.

Input
  │
  ▼
┌─────────────┐
│  Executor   │  ← ReAct pattern (Thought → Action → Observation)
│  (Groq LLM) │  ← Tools: web search, calculator, yfinance, python repl
└──────┬──────┘
       │
       ▼
┌─────────────┐
│   Critic    │  ← Scores: factuality, completeness, clarity, task_alignment
│  (score/10) │  ← Weighted overall score
└──────┬──────┘
       │
       ▼
┌─────────────┐     score ≥ 7.5  ──▶ ┌──────┐
│   Router    │ ─────────────────────▶│ Meta │──▶ END
│             │     max iterations ──▶│      │
│             │     budget exceeded ─▶│      │
└──────┬──────┘                       └──────┘
       │ score < threshold
       ▼
┌─────────────┐
│   Refiner   │  ← Improves output using specific critique reasoning
└──────┬──────┘
       │
       └────────────────────────────▶ Executor (next iteration)

Features

  • Self-improving loop — executor → critic → refiner cycle with configurable convergence threshold
  • ReAct executor — Thought/Action/Observation reasoning pattern with real tool use
  • Structured critic — 4-axis rubric scoring with few-shot examples and score anchors
  • Loop memory — ChromaDB stores task strategies across runs
  • Circuit breaker — stops early if scores decline for 2 consecutive iterations
  • Token budget guard — stops gracefully (not a crash) if a task runs away on tokens
  • LangFuse tracing — optional; per-iteration spans with scores, latency, and token counts (no-ops if unset)

Tech Stack

Layer Technology
Orchestration LangGraph
LLM (primary) Groq — llama-3.1-8b-instant
LLM (fallback) HuggingFace — Meta-Llama-3.1-8B-Instruct
Tools Tavily (search), yfinance, calculator, sandboxed Python REPL
Vector memory ChromaDB
Observability LangFuse (optional)

Project Structure

loopforge/
├── main.py                    # CLI entry point
├── requirements.txt
├── core/
│   ├── graph.py               # LangGraph state graph
│   ├── state.py               # GraphState TypedDict
│   ├── router.py              # Conditional edge logic
│   └── nodes/
│       ├── executor.py        # ReAct agent with tool use
│       ├── critic.py          # Rubric scorer (4 axes)
│       ├── refiner.py         # Critique-driven improver
│       └── meta.py            # Strategy memory + status resolution
├── tools/
│   ├── search.py               # Tavily web search
│   ├── calculator.py           # AST-safe math evaluator
│   ├── python_repl.py          # Sandboxed Python executor
│   └── yfinance_tool.py        # Market data
├── memory/
│   └── chroma.py                # ChromaDB client
└── observability/
    └── langfuse_client.py        # Optional trace + span management

Quickstart

Prerequisites

  • Python 3.11
  • API keys (see below)

1. Clone and configure

git clone https://github.com/Sahojit/Loop-Forge.git
cd Loop-Forge
cp .env.example .env

Edit .env and fill in the keys you need (at minimum GROQ_API_KEY; TAVILY_API_KEY if you want the executor to use web search):

GROQ_API_KEY=gsk_...             # console.groq.com/keys
HUGGINGFACE_API_KEY=hf_...       # huggingface.co/settings/tokens (fallback LLM)
TAVILY_API_KEY=tvly-...          # app.tavily.com (search tool)

2. Create virtualenv and install dependencies

python3.11 -m venv .venv
source .venv/bin/activate        # Windows: .venv\Scripts\activate
pip install -r requirements.txt

3. Run a task

python3 main.py "explain quantum entanglement in simple terms"

Or run it with no argument to be prompted for a task:

python3 main.py

Output includes the final answer plus the run's status, iteration count, score history, and tokens used.


Configuration

Variable Default Description
CONVERGENCE_THRESHOLD 7.5 Minimum critic score to stop the loop
MAX_ITERATIONS_DEFAULT 5 Max executor→critic→refiner cycles per task
TOKEN_BUDGET_PER_TASK 8000 Max tokens before the run stops early (status: budget_exceeded)
CHROMA_PERSIST_DIR ./chroma_db Where strategy memory is stored
LANGFUSE_PUBLIC_KEY / LANGFUSE_SECRET_KEY / LANGFUSE_HOST unset Optional tracing — no-ops if not set

How the loop ends

Each run finishes in one of these states, visible in the CLI output as status:

Status Meaning
converged Critic score reached CONVERGENCE_THRESHOLD
max_iter_reached Hit max_iterations without converging, or scores declined for 2 consecutive iterations (circuit breaker)
budget_exceeded Ran out of TOKEN_BUDGET_PER_TASK mid-run; returns the best partial output so far

License

MIT

About

Production-grade self-improving agent loop (LangGraph) with ReAct execution, rubric-based critique, and iterative refinement — FastAPI, Celery, Postgres/pgvector, JWT auth, LangFuse tracing.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages