Skip to content

Latest commit

 

History

113 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

attobot

A persistent agent in a single agent.py.

One agent = one working directory (in production, one unix user's $HOME) + one process running agent.py. All agent state lives in ./agent/. The process is a loop that re-runs the LLM every time messages.jsonl changes.

The loop

hash messages.jsonl
  if unchanged: sleep
  build system prompt + sum chars
  if over budget: stash_messages
  llm(<soul> + <harness> + <memory>, messages + [<life-tail>], tools)
  append assistant reply
  if tool_calls: run each via bg_run, append results
  else: automatic telegram text reply

That's agent.py. Channels, tools, and the backgrounding wrapper all live inline in the same file.

Channels in

Three daemon threads append to messages.jsonl:

  • telegram — start_chat() long-polls getUpdates. Inbound text → {role:user, content:"[telegram <id>] …"}. Only the chat_id/thread_id locked in config.json is accepted. Optional: no telegram_token in config → no chat channel; the agent wakes on triggers/mail only.
  • triggers — start_triggers() scans agent/triggers/*.json every 30s; due ones append {role:user, content:"<system-message>[trigger <name>] …</system-message>"}. Three kinds: a cron — {"next": <ts>, "repeat_s": <s?>, "message": "…"} — fires on the clock (repeating ones reschedule, one-shots delete); a watch — {"watch": "<path>", "repeat_s": <cooldown?>, "message": "…"} — fires when the file's content changes; a cmd — {"cmd": "<shell>", "repeat_s": <s?>} — runs the command (60s timeout) and fires with its stdout (clipped), no output = no fire. Combined with watch, the cmd runs when the file changes but receives no stdin; commands that need context should read files themselves. Fires are queued and injected one at a time, only when the stream is idle (last line is a plain assistant reply, no tool_calls in flight); if the previous trigger got only an idle text reply, that pair is collapsed before the next one lands, so unactioned triggers don't pile up. Triggers named subconscious-* are additionally surfaced to the operator's Telegram.
  • mail — start_inbox() polls agent/mail_inbox/. New files append {role:user, content:"[mail from <unix-user>] <name>\n<preview>"} and notify the operator via chat.

Fired triggers are persisted in agent/trigger-queue/ before one-shot schedules are removed. Conversation rewrites use a private recovery journal alongside messages.jsonl, keeping the existing file lock and inode; the next harness read or write completes an interrupted rewrite before proceeding. Delivery is at-least-once across interruptions, not an exactly-once guarantee for tool actions.

There is no mid-stream role:system — the only system message in a request is the system prompt itself. System-ish injections (triggers, bg completions, stash markers, the start banner) are user messages wrapped in <system-message>…</system-message>; the harness strips the wrapper when checking prefixes.

Channel out

Text replies are sent to telegram automatically — except when the turn's inbound was a <system-message>-wrapped injection (trigger, bg completion, stash marker) and the turn did no tool work: those replies are muted (this is what makes an idle heartbeat reply free). SEND_ATTACHMENT sends files to telegram and rejects text-only sends.

Tools

Declared in the TOOLS list in agent.py: (NAME, fn, description, parameters) per entry.

name what
SEND_ATTACHMENT send a file to telegram (photo/voice/video/audio/document by extension; optional caption)
READ_FILE file → line-numbered text; images → multimodal content blocks (when multimodal_support=true)
WRITE_FILE / EDIT_FILE filesystem writes; EDIT_FILE has optional replace_all
BASH run a shell command (returns Popen → streamed by bg_run)
SEARCH / WEB_FETCH DuckDuckGo + markdownified page fetch
STASH content-addressed save to agent/blobs/<hash>, returns [stash <hash>]
STASH_MESSAGES collapse a line range of a messages.jsonl (default: own, middle half) into one [stash <hash>] line with an LLM summary

Every tool call runs through bg_run; a tool returning a subprocess.Popen (or anything with .pid + .communicate) is streamed through it.

Tool results longer than tool_output_limit (5000 chars) are auto-clipped to <head>\n... N chars truncated, [stash <hash>] ...\n<tail>. The agent recovers the full content with READ_FILE agent/blobs/<hash>.

Backgrounding

bg_run runs the tool in a thread with tool_timeout (30s). Finishes in time → inline (post-clip) result. Otherwise the work keeps running in the background:

  • registers agent/bg/<id>.json (with pid if known)
  • returns [backgrounded bg/<id> (pid …); kill the pid to stop it] to the assistant immediately
  • an emitter thread waits for the worker to finish — however long that takes — then appends [bg <id> done, tc:…] <result> as a <system-message> user message and removes the json

The bg json holds the pid so the agent can kill a runaway task itself. Background work does not survive a process restart (the worker is a daemon thread); anything that must outlive the harness should detach itself (e.g. nohup … &).

State

SOUL.md                       # the prompt template for agent/SOUL.md
agent.py                      # the harness, included verbatim in the system prompt
opt/
  tools/<name>.py             # optional capability tools (see Optional add-ons)
  providers/<name>.py         # alternative LLM providers
  subconscious/               # reviewer agent skeleton (soul + seeded trigger), copied out beside agent/
agent/
  SOUL.md                     # this agent's soul (copy of the template)
  MEMORY.md                   # memory index: one pointer line per memory
  memory/<name>.md            # memory bodies, read on demand via the index
  LIFE.md                     # append-only event log; tail rides as a trailing user message
  messages.jsonl              # canonical conversation, one JSON message per line
  config.json                 # telegram token/chat, api key, overrides
  tg_poll.offset              # telegram update_id cursor
  triggers/<name>.json        # crons (clock), watches (file change), cmds (computed)
  triggers/heartbeat.json     # auto-created at boot (not for subconscious dirs), 225s tick, backs off when idle up to 1h
  mail_inbox/                 # drop files here (delivered once, then moved to processed/)
  inbound/                    # files received over telegram
  bg/<id>.json                # in-flight background work
  tools/<name>.py             # opt-in tools (copied from opt/tools/ at first boot)
  providers/<name>.py         # opt-in provider (copied from opt/providers/ at first boot)
  blobs/<hash>                # content-addressed store

System prompt

Built fresh every turn:

<soul>          agent/SOUL.md
<harness>       agent.py source
<subconscious>  one-liner, present only when a subconscious/ sibling dir exists
<memory>        MEMORY.md (middle-elided if > MEMORY_LIMIT)

The last life_tail lines of LIFE.md (prefixed with [N bytes earlier]) are not part of the system prompt — they ride along as a trailing <system-message> user message after the conversation, rebuilt every turn.

The agent sees its own harness. The source is memoized at first read — modify agent.py and the self-model updates on restart.

Memory pressure

  • Memory is two-tier: MEMORY.md is an always-in-context index (one pointer line per memory), bodies live in agent/memory/ and are read on demand. MEMORY.md > MEMORY_LIMIT (10000) → middle is elided with a warning telling the agent to move detail into agent/memory/ files.
  • System prompt + life-tail + serialized messages, divided by 4 chars/token, > context_tokens * 0.8 → stash_messages runs automatically; the middle half of messages.jsonl goes to a blob, replaced by a single <system-message>-wrapped user message holding [stash <hash>] plus an LLM summary (ranges snap past tool messages so tool-call blocks stay intact). The agent can READ_FILE agent/blobs/<hash> to recover.
  • A user message whose content is STASH_MESSAGE: <start> <end> or STASH_MESSAGE: all (bare, or right after a [trigger …] prefix) is a stash directive: the loop intercepts it before calling the LLM, removes the directive line, stashes the range, and owes the agent a re-orientation turn. This is what the subconscious's PRUNE rides.

The 4-chars-per-token heuristic over-counts base64 image content — safe direction. load_messages() also heals the file on every read: malformed lines and incomplete tool-call blocks are dropped and the file is rewritten.

Optional add-ons

Anything under opt/ is opt-in via the opt field in config.json (a list of paths relative to opt/, no .py suffix):

"opt": ["tools/ocr_image", "providers/anthropic"]

Each entry copies opt/<path>.py → agent/<path>.py at first boot. From then on, the agent owns its copy.

Tools in agent/tools/ auto-register at startup. Built-in:

  • ocr_image — RapidOCR + spatial ASCII layout, for text-only LLMs. Auto-included when multimodal_support=false. Requires rapidocr-onnxruntime + opencv-python.
  • nudge (NUDGE) / stash_messages (PRUNE) — the subconscious's correction tools (see Subconscious); they act on the sibling agent/ dir, so they only make sense in a subconscious's opt list.

Providers swap the llm function. Set provider: "anthropic" (auto-includes providers/anthropic) to use it. Built-in:

  • anthropic — native /v1/messages translation. Reads api_key and model from config.json like the default provider.

Subconscious

A second attobot that reviews the first. Same harness, different soul (opt/subconscious/ — an agent-dir skeleton: soul + a pre-seeded watch job), no chat. It runs in the same unix user as the primary — it needs direct read/write into agent/ — unlike peer agents, which get a user each.

The sibling directory needs its own config.json with the provider settings and "opt": ["tools/nudge", "tools/stash_messages"]. These entries install the tools required by its soul. The supplied triggers assume agent/ and subconscious/ under the working directory; custom layouts need corresponding trigger paths and primary_dir in the subconscious config. Attosys provisions the standard layout.

Run both configured directories with:

python agent.py agent subconscious

A pre-seeded selfwipe.json trigger stashes the subconscious's own messages.jsonl to a single pointer every ~30min (only when it has grown past 20 lines) — the self-wipe its SOUL describes.

It wakes via its pre-seeded primary.json trigger — a watch+cmd on agent/messages.jsonl with a 450s cooldown whose cmd diffs the stream against a snapshot (subconscious/.primary_snap) and fires with the new lines as the trigger message, but only when more than 10 lines are new (no heartbeat: the harness skips heartbeat creation for dirs named subconscious). It corrects the primary via two opt tools (opt/tools/nudge.py, opt/tools/stash_messages.py — copy them into subconscious/tools/, where they register as NUDGE and PRUNE):

  • NUDGE — writes a one-shot trigger agent/triggers/subconscious-<name>.json; the primary's trigger thread fires it as a [trigger subconscious-<name>] <message> injection on its next tick and, because of the subconscious- prefix, also surfaces it to the operator's Telegram.
  • PRUNE — writes a one-shot trigger carrying a STASH_MESSAGE: <start> <end> directive; the primary's loop intercepts it and collapses that line range of messages.jsonl into a stashed summary pointer. Use on context rot or to refocus the primary.

Both act through the trigger-file bus — the subconscious never writes the primary's messages.jsonl directly, so the stream can't be corrupted.

Run

Attobot runs a configured agent directory. Attosys handles company provisioning, Unix accounts, services, and deployment.

Before starting, the state directory must contain config.json (a JSON object) and SOUL.md (use the root template or your own prompt). Supply the LLM key through ATTOBOT_API_KEY or the config's api_key field; the environment takes precedence. Keep configurations containing credentials private (mode 600).

pip install -r requirements.txt
python agent.py agent

python agent.py [agent_dir ...] defaults to ./agent. Extra directories each get their own process; Ctrl-C stops them all. A child that dies is respawned after 10s and the death is logged to the primary's LIFE.

The default model is deepseek-v4-pro. Override model / api_base in the config for an OpenAI-compatible endpoint, or select an adapter with provider.

Omit telegram_token for a chat-less agent. For Telegram, configure the token and chat ID, plus a thread ID if needed. The bot must be allowed to join groups and have privacy mode disabled to receive all group messages.

agent/config.json fields (defaults live in agent.py):

{
  "telegram_token": "...",         // optional — omit for no chat channel
  "telegram_chat_id": "...",       // required if telegram_token is set
  "telegram_thread_id": "...",     // optional, forum supergroup topic
  "api_key": "...",                // optional when ATTOBOT_API_KEY is set
  "model": "deepseek-v4-pro",
  "api_base": "https://api.deepseek.com/v1",
  "temperature": 1.0,
  "reasoning_effort": "medium",
  "context_tokens": 100000,
  "multimodal_support": false,
  "provider": "",                  // "" = openai-compat default; "anthropic" loads opt/providers/anthropic
  "opt": []                        // additional opt/ entries to copy in
}

Tunables with defaults in CFG (rarely worth changing, override in config.json): life_tail, memory_limit, tool_timeout, trigger_tick, inbox_tick, inbox_preview, chat_msg_max, tool_output_limit. AGENT_DIR / BLOB_DIR are in-source constants.

The Responses provider allows at most three generation attempts, doubling the per-request budget up to max_tokens_limit (default 32768, or the initial budget if higher). It never executes incomplete tool calls. If that recovery is exhausted, the employee stays alive, logs the limit, and waits for new input instead of retrying indefinitely or exiting with a permanent configuration error.

Verification

Run the local regression suite from this checkout:

python3 -m unittest test_agent test_responses -v

It covers trigger persistence, interrupted conversation rewrite recovery, concurrent message writers, startup validation, process locking, web fetching, and provider output-limit recovery. It makes no live model calls.

Principles

  1. The agent is a loop. One process, one file watch, one LLM call per change.
  2. The bus is the filesystem. Channels in, channels out, scheduled jobs, background work, memory — all files. No daemon, no queue, no IPC.
  3. Opinionated cuts code. Telegram is the chat. One operator, one chat. Default is DeepSeek V4 Pro via DeepSeek, but anything OpenAI-shape works out of the box and other shapes live in opt/providers/. No abstractions for things that aren't pluralized.

About

very smol agent harness

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages