Boku documentation

Boku is an open-source multi-agent research engine written in Go. It plans a research programme, runs specialised agents in parallel, keeps every claim linked to its sources, has an independent critic check the evidence, and renders a professional report as PDF, HTML and Markdown.

Agents supply content and data; Boku's own code does everything else — de-duplicating sources, resolving citations, checking that chart values appear in the evidence, and typesetting. By default agents run on Claude Code. With a config file they run on any OpenAI-compatible endpoint or a local Ollama model.

Install

Requirements:

go install github.com/riteshsonawane1372/boku/cmd/boku@latest
boku doctor

or from source:

git clone https://github.com/riteshsonawane1372/boku
cd boku
make build        # ./bin/boku

boku doctor checks each piece:

  ✓ config       ok
  ✓ claude code  2.1.285 (Claude Code)
  ✓ ollama       llama3.1:8b at http://localhost:11434 (--quick and local formatting available)
  ✓ chrome       /Applications/Google Chrome.app/Contents/MacOS/Google Chrome
  ✓ prompts      ok

Quick start

# Full report (default)
boku report "How enterprises are deploying AI agents in 2026"

# A short search report: a cited brief of a few pages
boku report "Is Postgres 18 async I/O worth enabling?" --short

# Quick report on your local Ollama model: no Claude Code tokens
boku report "Main approaches to GPU sharing on Kubernetes" --quick

# Put the numbered source list inside the PDF too
boku report "…" --save-ref

# Your own model instead of Claude Code
boku report "…" --config config.yaml

Every run writes a directory under runs/; published files go to reports/. Interrupted runs continue with boku resume <run-dir>.

Report modes

You never need a flag for the default. Pick another mode per report:

ModeFlagWhat it does
Fullnone (or --mode full)Planner, up to 6 research workstreams, fact-check with follow-up research rounds, synthesis, editorial revision. Full layout: cover, contents, sections, methodology.
Short--short (--mode short)Quick depth, at most 3 workstreams, one fact-check round without follow-ups, and a brief of 2–4 sections in the compact layout (title block, no cover or contents).
Quick--quick (--mode quick)Every role on the local Ollama model; Boku does the web search itself; fact-checking is skipped; compact layout. Spends no Claude Code tokens.
Whitepaperboku whitepaper, --whitepaper (--mode whitepaper)The most detailed format: a research paper with abstract, keywords, numbered sections and subsections, captioned figures and tables, [n] citations and a reference list. Deep research; references always included; paper layout. See whitepapers.
Explainerboku explain, --explainer (--mode explainer)A visual explainer instead of a report: big picture, key ideas, architecture and process diagrams, step-by-step flows, glossary. Up to 4 workstreams, one follow-up round. Given a directory, explains that codebase. See explainers.

Modes set defaults; explicit flags still win. --short --agents 4 keeps four workstreams, --quick --local-model qwen2.5:14b uses a bigger local model.

Quick reports are a first pass. A small local model reads fewer pages and nothing checks its claims. Boku still refuses sources it did not fetch and drops citations to unknown findings, and the methodology appendix says the report was not fact-checked. Use --short or the default when accuracy matters.

Whitepapers

boku whitepaper writes the subject up the way a conference paper or arXiv preprint would: a title block, an Abstract with keywords, optional Highlights, then numbered sections — Introduction (ending with the paper's contributions and organisation), Background, Related Work, the core technical sections, Empirical Evidence, Discussion, Open Problems and Future Directions, Conclusion — followed by References and lettered appendices (methodology, evidence register).

boku whitepaper "Sparse attention for long-context transformers"
boku report "Disaggregated prefill and decode for LLM serving" --whitepaper
boku whitepaper "…" --depth standard --max-cost 25     # cheaper; flags still win
First page of a whitepaper in Boku's paper layout: centred title between rules, author line, abstract, keywords, highlights and section 1
Layout preview: first page of the paper layout, rendered from Boku's fictional test fixture.

Explainers

boku explain runs the same pipeline — plan, research, fact-check, synthesis, editing — but the goal is understanding, not a verdict. The editor writes for a smart newcomer: The big picture first, then Key ideas, an architecture diagram of the whole, a section per main flow with a process diagram and numbered steps, a glossary, and Where to go next.

boku explain "How does Raft consensus work?"
boku explain ./path/to/repo                                  # explain a codebase
boku explain ./path/to/repo "how a request is authenticated" # …with a focus
boku report "How TLS 1.3 works" --explainer                  # same as boku explain

Codebases. When the first argument is a directory, the subject is that repository (research.codebase). The planner receives its file tree, README and current commit. The planner, researchers and fact checker get the read-only Read, Glob and Grep tools and run inside the repository; they still have no shell and cannot write. Sources are files, cited by repository-relative path and line range (internal/agent/claude.go#L48-L75), and the explainer adds a Where things live table mapping concepts to directories and files. The editor never sees the repository, only the checked findings.

Codebase explainers need the Claude Code provider, because other providers have no file tools, and cannot run with --quick. Agents can read any file in the directory you give them, so don't point Boku at a directory holding secrets you wouldn't send to the model provider.

References file

Citations in the report are numbers: … grew 34% in 2025.3. The sources behind those numbers are always written next to the report as compact JSON, one entry per line:

reports/
  kubernetes-platform-ai-infrastructure.pdf
  kubernetes-platform-ai-infrastructure.references.json
{"title":"Kubernetes as the Platform for AI Infrastructure","topic":"…","run_id":"…","date":"2026-09-23",
"references":[
{"n":1,"title":"…","publisher":"CNCF","url":"https://…","published":"2026-01-20","accessed":"2026-09-23","tier":1},
…
],
"evidence":[
{"id":"F052","claim":"Gartner forecasts …","status":"Estimated","as_of":"2026","refs":[5]},
…
]}

n is the citation number in the text; evidence lists every cited finding with its status and the references that support it. The PDF, HTML and Markdown leave the source list and evidence register out and say which file holds them — keeping the document short and making the references cheap to hand to another model.

To print them in the document as well, pass --save-ref (or set report.include_references: true). It works on re-renders too, at no cost:

boku render runs/2026-09-23T024958-kubernetes-used-ai-infrastructure --save-ref

Report shape follows the request

Reports are not poured into one template. The planner writes a report_shape from your wording, and the synthesizer and editor receive your original request with an instruction to honour its format, length, audience and structure:

The editor may rename the fixed parts (summary_title, key_findings_title, conclusion_title), leave out key findings or the conclusion when the form does not need them, and choose layout: compact. Set report.layout: full or compact to override its choice.

Web interface

boku ui runs the same pipeline behind a local web app. It is embedded in the binary: no extra install and no build step.

boku ui                          # serves http://127.0.0.1:7878 and opens the browser
boku ui --port 9000 --no-open
boku ui --config team.yaml       # read and save this file instead of ./boku.yaml

The New report page: mode cards, the settings form and a summary of the run

PageWhat it does
New reportA topic, or a local codebase and focus for explainers; the five modes; every setting in boku.yaml as a form. A side panel shows what the run will do, validates as you type, lists the settings you changed, saves presets, and prints the equivalent boku command.
RunsEvery run in the runs directory. A run page shows live stages, cost by stage, quality gates, agent calls and the log, with tabs for the report, plan, evidence, fact-check rounds, files and settings. Cancel, resume, re-render, duplicate or delete from there.
ReportsPublished PDFs, HTML, Markdown and reference files, each linked to its run.
SettingsEdit the defaults and save them to boku.yaml; the boku doctor checks.

Runs are ordinary run directories: runs started from the CLI appear in the UI, live, and boku resume continues runs started in the UI.

Stopping boku ui cancels the runs it is executing; resume them afterwards. Saving settings rewrites boku.yaml, so comments in an existing file are lost.

FlagMeaning
--port Nport to listen on (default 7878)
--host ADDRaddress to listen on (default 127.0.0.1). Anything but loopback requires the access token printed in the URL at start
--config FILEconfiguration file to read and save
--no-opendo not open the browser

The server listens on loopback only and rejects requests from other websites. Exposed with --host, it speaks plain HTTP behind a token: anyone with the URL can start agents and read run files, so keep it to networks you trust.

Commands and flags

boku report <topic> [flags]     research a topic and publish a report
boku explain <topic|dir> [focus] [flags]
                                visual explainer of a topic, or of a local codebase
boku whitepaper <topic> [flags] detailed paper in academic/industry format
boku resume <run-dir> [flags]   continue an interrupted or blocked run
boku render <run-dir> [flags]   rebuild outputs from a run's artifacts (no agents, no cost)
boku status <run-dir>           stages, tasks and cost of a run
boku ui [flags]                 open the web interface
boku doctor                     check the agent runtime, Ollama and Chrome
boku init [path]                write a commented boku.yaml
FlagMeaning
--short, --quick, --explainer, --whitepaper, --mode Mreport mode (see modes)
--save-refinclude the source list and evidence register in the document
--local-model TAGOllama model for --quick and formatting
--depth quick|standard|deephow much research to do
--agents Nmaximum research workstreams
--parallel Nmaximum agents running at once
--max-cost USDstop launching agents at this spend
--freshness 30d|6m|1ywindow within which information counts as current
--sources "…"preferred sources or source types
--min-sources Ndistinct sources the research gate expects
--format pdf,html,mdoutputs to publish
--model Mmodel for all agents on the main provider
--iterations Nfact-check → follow-up research rounds
--config FILEconfiguration file (default ./boku.yaml if present)

Exit codes: 0 published, 1 error, 3 blocked by a quality gate, 130 interrupted.

Run directory

runs/2026-09-23T074500-enterprises-deploying-ai-agents-2026/
  manifest.json      configuration, stage status, every agent call, cost
  plan.json / .md    planner output
  research/          one JSON artifact per research task
  evidence/          de-duplicated sources and findings (F001 ⇄ S001)
  factcheck/         one artifact per critic round
  synthesis/         thesis, insights, outline
  report/            document.json (drafts), report.json, references.json, gates.json, HTML, Markdown
  output/            the PDF before publishing
  agents/            raw agent responses
  logs/boku.log

Custom models

Claude Code is the default and needs no configuration. To run Boku on another model, pass a config file. Nothing changes unless you do.

# config.yaml
agents:
  provider: openai               # claude-code | ollama | openai
  endpoint: http://localhost:8000/v1
  model: Qwen/Qwen2.5-72B-Instruct
  api_key_env: OPENROUTER_API_KEY # name of the env var; empty for a local server
  price_input_per_mtok: 0.35     # optional, for cost tracking and --max-cost
  price_output_per_mtok: 0.40
  timeout: 30m
boku report "How are banks deploying generative AI?" --config config.yaml
ProviderSpeaksExamples
claude-codeclaude -p in restricted modedefault; agents.model: opus, role_models
openaiPOST {endpoint}/chat/completionsvLLM, SGLang, LM Studio, llama.cpp server, OpenRouter, OpenAI, Together, Groq
ollamaPOST {endpoint}/api/chatany pulled Ollama model

Structured output: Boku sends the JSON Schema for each agent's output as response_format: json_schema (OpenAI-compatible) or format (Ollama), and also in the system prompt. Servers that reject response_format are retried without it.

Mix providers per role with local.roles: for example keep research on Claude Code and run synthesis and formatting locally with local.roles: [synthesizer, formatter]. See examples/custom-model.yaml and examples/local-ollama.yaml.

Local model (Ollama)

ollama pull llama3.1:8b        # or qwen2.5:7b, qwen2.5:14b, gemma2:9b …
boku doctor

The local model does two jobs:

  1. Formatting pass (always, when Ollama is reachable). If the editorial draft passes every check but has style problems — generic phrasing, stray Markdown, repeated sentences — Boku sends just those passages to the local model instead of asking the main model for a full revision. A rewrite is accepted only if it keeps exactly the same citation markers and numbers, so it cannot change facts. Turn it off with local.format: false.
  2. Quick mode. --quick routes every role to the local model.
local:
  endpoint: http://localhost:11434
  model: llama3.1:8b
  context_tokens: 16384          # num_ctx; Ollama's default of 2048 is too small
  format: true
  roles: []                      # e.g. [formatter, synthesizer]; --quick uses all

Boku fits fetched pages to context_tokens (Ollama silently drops the start of an over-long prompt) and caps each generation, so a small model cannot run away. Larger models (14B+) with 32k context write noticeably better quick reports.

Claude Code searches and reads the web by itself. Ollama and OpenAI-compatible models cannot, so Boku does it for them before each research task:

  1. Search each workstream question (DuckDuckGo by default, or your SearXNG instance).
  2. Fetch the top result pages, round-robin across questions, and extract readable text, title, publisher and publication date.
  3. Give the pages to the model as numbered sources s1…sN.
  4. Re-ground the output: a cited source must be a page Boku fetched; its URL, title and date come from the fetch, not the model. Findings citing anything else are rejected as unsourced.
search:
  engine: duckduckgo             # duckduckgo | searxng
  searxng_url: ""                # e.g. http://localhost:8888 (JSON output enabled)
  results_per_query: 5
  max_pages: 10                  # per research task
  page_chars: 4000               # text kept per page

Page fetches refuse loopback, private and link-local addresses: result URLs are untrusted, so a fetch can never reach services on your machine or network.

Configuration

boku init writes a commented boku.yaml. Boku reads ./boku.yaml automatically, or the file given with --config. Flags override the file; unknown keys are errors.

agents:
  provider: claude-code
  model: ""                      # Claude Code default
  role_models: {editorial: opus}
  max_parallel: 4
  max_agents: 6
  max_retries: 2
  timeout: 20m
  max_cost_usd: 0                # 0 = unlimited

local:  {endpoint: http://localhost:11434, model: llama3.1:8b, context_tokens: 16384, format: true}
search: {engine: duckduckgo, results_per_query: 5, max_pages: 10, page_chars: 4000}

research:
  depth: standard                # quick | standard | deep
  freshness_days: 365
  max_iterations: 2
  min_sources: 0                 # 0 = from depth: 8 / 20 / 35
  fact_check: true
  codebase: ""                   # set by boku explain <dir>

report:
  mode: full                     # full | short | quick | explainer | whitepaper
  layout: auto                   # auto | full | compact | paper
  include_references: false      # --save-ref
  formats: [pdf]
  page_size: A4
  author: ""

output:
  directory: ./reports
  runs_directory: ./runs
  prompts_directory: ""

Pipeline and agents

topic ─▶ Planner ─▶ plan.json (objective, questions, workstreams, report shape)
            ├─▶ Primary ─┐
            ├─▶ Market ──┤  parallel research agents
            ├─▶ Technical┤
            └─▶ …       ─┘
                          ▼
                 Evidence store  (F001… ⇄ S001…, de-duplicated, tiered, dated)
                          ▼
                 Fact checker ⟲ follow-up research      ── research & fact gates
                          ▼
                 Synthesis (thesis, insights, outline)
                          ▼
                 Editorial ⟲ local formatter ⟲ revision ── editorial gate
                          ▼
                 Report model ─▶ HTML ─▶ PDF             ── PDF gate
                              └▶ references.json
RoleJobTools
plannerobjective, questions, workstreams, report shapeweb search
primary · market · technical · financial · competitive · case-studyresearch one workstream into sourced findingsweb search, fetch
fact-checkerverifies claims, rejects unsupported ones, requests follow-upsweb search, fetch
synthesizerselects, connects and interprets; designs the outlinenone
editorialwrites the report as structured content with citationsnone
formatterfixes style problems in flagged passages (local model)none

Claude Code agents run as claude -p in restricted mode with only web tools (plus read-only file tools for codebase explainers), no MCP servers, a filtered environment and a JSON Schema for their output. The Go process orchestrates; agents never write files or run commands.

Source breadth

Researchers are told to go wide before going deep: search each question several ways, mix source types (primary documents, independent research, quality press, practitioner write-ups), follow citations upstream to the original, corroborate headline numbers with two independent sources, and look specifically for the last 90 days. Targets by depth:

DepthFindings per researcherSources per researcherResearch gate target
quick6–106+8 sources
standard12–2010+, three types20 sources
deep20–3518+, four types35 sources

The planner names 4–8 specific priority sources per workstream. Raise the bar with --min-sources.

Quality gates

A blocked report exits with code 3; the reason is in report/gates.json and nothing is published.

Cost

Costs reported by Claude Code for real runs, recorded in each run's manifest.json:

RunDepthAgent callsFindings / sourcesCost
Kubernetes for AI infrastructurequick1172 / 77$7.73
GPU monitoring tool roadmapstandard13120 / 147$12.96
Self-hosting vs managed AIdeep10197 / 217$24.09
Launching an agentic AI start-updeep10178 / 207$30.72

Research agents account for 70–85% of spend. To spend less: --short, --agents 3, --iterations 0, a cheaper model for researchers via role_models, a hard cap with --max-cost, or --quick for zero API spend.

Tuning

Troubleshooting

SymptomFix
quick mode needs the local modelStart Ollama and ollama pull llama3.1:8b, or pass --local-model with a model you have.
duckduckgo returned HTTP 202 / no resultsYou are being rate-limited; lower agents.max_parallel or use search.engine: searxng.
Exit code 3A gate blocked publication; read report/gates.json. Resume after raising limits, or re-render.
No PDFInstall Chrome/Chromium or set report.chrome; or --format html,md.