Blog
Building in public. Sharing what we learn.

Octofs 0.9.0: Line 42 Is a Lie
Octofs 0.9.0 gives every line a content-verified composite ID (N:hh) that edit tools check before touching a file, so a stale target fails loudly with the fresh content instead of silently editing the wrong line. Plus replace_all and CRLF-safe matching in str_replace, and hard rejection of shell misuse. Open source, Apache-2.0.

Introducing Pulsora: A Rust Time-Series Database for Trading Data
We built Pulsora, an open-source time-series database in Rust, because a firehose of market ticks pushed our storage past its limits. Columnar blocks, type-specific compression, and a block-level index that doesn't grow per row.

Introducing Synx: File Sync for Remote Development
Synx is an open-source, real-time two-way directory sync over SSH — a Mutagen alternative for developers who edit locally and build on a remote box. Written in Rust, one command, no daemons.

Introducing OctoHub: One Front Door for Every LLM Your Agents Call
OctoHub is a self-hosted LLM proxy you run in front of your agents. One endpoint, model aliases, per-provider load balancing, and a log of every request that went out and what it cost. Open source, Rust, Apache-2.0.

Octocode 0.20.0: Markdown Joins the Knowledge Graph
Octocode 0.20.0 brings Markdown into the GraphRAG knowledge graph — docs become nodes, cross-document links become typed references relationships — plus a fix for silently missing relationships on large projects and a 2x memory cut on relationship loading. Open source, Apache-2.0.

Debugging AI Agents: The Observability We Wished We Had on Day One
An agent burned $4 and 90 tool calls on a task that should have taken three, and we were staring at a blank terminal. This is the instrumentation that turned guesswork into a two-minute diagnosis — Octomind's /info, /report, /context, the zstd session log, RUST_LOG tracing, --format jsonl, and OctoHub in front to capture every upstream request.

Release Round, Late July 2026: Octocode 0.19.0, Octobrain 0.9.4, Octolib 0.26.1, Vext 1.3.0
Two and a half weeks since the July round, and the stack learned to check its own work. Octocode 0.19.0 shipped reasoning retrieval — an LLM re-ranker fused into hybrid search with weighted RRF, +36% MRR on the benchmark. Octobrain made knowledge sync non-blocking, Octolib kept the model roster and embedding pricing current, and Vext 1.3.0 landed a glass redesign, word-level diarization, and two new languages.

A Map and a Memory: Pairing Code Search with Persistent Memory for AI Agents
Semantic code search lets an agent find the right code. Persistent memory lets it remember the decisions about that code. Run only one and you get an agent that re-derives context every session or remembers conclusions it can't relocate. Here is how to wire Octocode and Octobrain together so the agent both finds and remembers.

Reasoning Retrieval: We Taught Code Search to Think, Not Just Match
Octocode now has an optional LLM reasoning step that reasons over retrieved code and re-ranks by real relevance, fused with hybrid search via RRF. On a 127-query benchmark it lifts MRR +36% and Hit@5 to 0.953 — with every metric up. Here are the numbers, the tuning, and what did not work.
One SQLite File, 1 Hz, Zero Cloud: The Local-First Architecture of a Mac Time Tracker
A grounded design guide to building a local-first desktop app: why one embedded SQLite file beats a cloud DB for a single user, how 1 Hz raw samples roll up into human sessions, schema design for time intervals, WAL and crash safety, and the honest tradeoffs of owning your own data.

Why We Run Speech Recognition Fully On-Device: The Latency and Privacy Math
Cloud speech-to-text loses on a budget you can compute on a napkin. Here is the round-trip math that forced Vext to run both the speech model and the cleanup LLM on-device, plus a cloud-vs-local decision framework for your own audio feature.

Lessons From Building a Unified LLM Provider Layer in Rust
What we learned shipping octolib, the Rust LLM library behind our AI stack: how to abstract multiple LLM providers, normalize tool calls and token usage, and survive the day a model's thinking format returned a 400.