dendrux
v0.2.0a1 · alphaGet started

How to run dendrux inside a stateless HTTP server — build the Agent per request, share the expensive resources, switch model per turn, and resume long runs across processes.

dendrux in a web or chat endpoint

dendrux is built to run inside your server. This recipe is the architecture for doing that well: what to build per request, what to share, how to switch model per turn, and how a single long-running task survives across requests.

The one idea

The Agent is a disposable executor. The run is the durable thing.

  • dendrux persists run state (runs, events, traces, tool calls, LLM calls, pauses) to your database.
  • It does not persist the Agent object or the conversation history. Your app owns the conversation.

Because state lives in the DB, you can build a fresh Agent on each request — even on a different machine — and resume any run by ID. Constructing an Agent does no I/O, so this is cheap. The Quickstart shows the proof: a run started in one process is resumed in another.

Build per request, share the expensive parts

Building the Agent object is essentially free. The cost is in the resources it connects to. Build those once per worker and reuse them:

ResourceCost if rebuilt per requestWhat to do
DB engineA new connection pool every requestShare one: set DENDRUX_DATABASE_URL or inject a state_store.
MCP sourceRe-spawns a subprocess / reconnects + re-lists toolsShare one process-wide MCPRuntime; bind() per request with the requester's tenant_key.
Provider clientLoses connection keep-aliveUse a request-owned provider unless your application has an explicit non-Agent owner for a shared provider.
# ---- once per worker (startup) ----
from dendrux.db.session import get_engine
from dendrux.mcp import MCPRuntime, MCPSource
from dendrux.runtime.state import SQLAlchemyStateStore
 
engine = await get_engine(os.environ["DENDRUX_DATABASE_URL"])  # shared singleton
STORE  = SQLAlchemyStateStore(engine)
MCP_RUNTIME = MCPRuntime()                                     # one per worker process
MCP_SOURCES = {
    "github": MCPSource.http(
        "github",
        os.environ["GITHUB_MCP_URL"],
    )
}
 
class GitHubCredentials:
    def __init__(self, user_id: str) -> None:
        self.user_id = user_id
 
    async def get_auth(self) -> dict[str, str]:
        # Your application owns encrypted storage, refresh, and authorization.
        token = await credential_store.github_access_token(self.user_id)
        return {"Authorization": f"Bearer {token}"}
 
# ---- per request ----
@app.post("/chat")
async def chat(req):
    views = [
        MCP_RUNTIME.bind(                       # synchronous, no I/O; cheap per request
            connection_key=m,
            tenant_key=req.user_id,             # per-tenant connections and credentials
            source=MCP_SOURCES[m],
            credentials=GitHubCredentials(req.user_id),
        ).tools()
        for m in req.enabled_mcps
    ]
    async with Agent(
        provider=f"{req.vendor}:{req.model}",             # request-owned provider
        prompt=req.system_prompt,
        tools=[TOOLS[t] for t in req.enabled_tools],          # current selection
        tool_sources=views,                                   # shared runtime leases
        state_store=STORE,                                    # shared engine
    ) as agent:
        result = await agent.run(
            req.text,
            history=req.transcript,                           # your app owns this
            metadata={"thread_id": req.thread_id},
        )
        return {"answer": result.answer}

Ownership at request shutdown

MCPRuntime makes MCP ownership explicit: agent.close() releases the agent's MCP leases but does not close shared MCP connections. Always close every request-scoped Agent (the async with above guarantees it), or its lease can pin a connection and delay runtime shutdown. The runtime keeps each tenant's connection warm for the next request, retires it after idle_timeout, and the application closes the runtime during worker shutdown. See MCP for tenancy, credentials, capacity, and telemetry.

An Agent owns and closes the provider passed to it, so this example uses a provider recipe string to create one per request. Do not hand an Agent a shared provider unless your application has deliberately coordinated that provider's ownership outside the Agent lifecycle. Close application-owned pools at worker shutdown:

@app.on_event("shutdown")
async def shutdown():
    await MCP_RUNTIME.close()   # drains active MCP work, closes every transport

Shortcut for simple scripts: a provider recipe string

For a one-off script (not a pooled server) you can skip constructing the provider yourself:

async with Agent(provider="anthropic:claude-haiku-4-5", prompt="...") as agent:
    print((await agent.run("hi")).answer)

provider="vendor:model" builds the provider for you — reading the API key from the environment — the same way database_url builds an engine, and async with closes it. Supported vendors: anthropic, openai, openai-responses. Pass a provider instance when you need full configuration (custom base_url, explicit api_key, sampling defaults) or when pooling across requests.

Switching model or vendor per turn

For a "pick your model per message" chatbot:

  • Model (same vendor): pass it to run(). No new Agent needed.

    await agent.run(req.text, history=req.transcript, model="claude-opus-4-1")

    The model actually used is recorded per call, so RunStore.get_llm_calls reflects the real model for billing and observability.

  • Vendor, prompt, or tool set: rebuild the Agent with the new provider / prompt / tools. Construction is cheap, and rebuilding is the clean way to reflect a capability set the user changed mid-chat.

New turn vs. continuing a long run

A long-running task usually pauses and resumes rather than finishing in one call. Use the right entry point:

  • New turnagent.run(...) starts a new run.
  • Continue a paused runagent.submit_tool_results() / submit_input() / submit_approval() / resume() on the existing run_id.

Both work with a freshly built Agent, possibly in a different process. One thing to know: on resume, dendrux replays the conversation from the DB but takes the provider, model, prompt, and tools from the Agent you resume with — not from a saved snapshot. So keep the config consistent across the resume. The clean way is to stash a config key when you start the run and read it back:

# start
await agent.run(text, metadata={"thread_id": tid, "config_id": cfg_id})
 
# later — resume with the SAME config the run started with
run = await store.get_run(run_id)
agent = build_agent(load_config(run.meta["config_id"]))
await agent.submit_tool_results(run_id, results)

New turns can use the user's latest model/tool selection; an in-flight run should be resumed with its own config (especially if it paused waiting on a client tool the user may have since toggled off).

Where this fits