Jarxi

Bolderdash

One API Client, Two Authentication Channels

The web uses an HTTP-only cookie, the CLI uses a bearer token. The server checks both in one middleware, and each shell injects its own way of producing credentials — so feature code never learns which platform it is on.

Following on from the previous post: three shells sharing one set of logic packages. What about auth? The web wants an HTTP-only cookie — sent automatically, unreadable from JS, safe against XSS. T......

Next.js Is Just a Shell, Peer to Electron

The easiest mistake in a multi-client monorepo is treating the Next.js app as the product and everything else as a port. Demote it to a shell and the package boundaries fall out — especially the one the CLI forces on you.

Say a product needs a web app, a desktop app and a CLI. The obvious path is to build the Next.js app first, then find a way to stuff it into Electron, then find something for the CLI to reuse. That......

Cross-Package Imports Are Type-Only, Runtime Goes Through ctx

import type compiles to nothing, so it creates no runtime dependency edge. That is how dozens of UI packages can know each other's types while never loading each other — the plugin boundary lives in one keyword.

I was reading an “everything is a plugin” frontend codebase — dozens of UI packages that almost never import each other. One package opened with these two lines: 12import type { Context }......

Judge Your Judge Twice Before You Trust It

I used a large model as a judge to fact-check SFT data. Judging the same note twice gave zero errors on one run and a dozen on the next. A single verdict is a coin flip.

I have been turning a batch of meeting transcripts into SFT data, so a small model can write meeting notes against a fixed template. After the pipeline was running and the filters were written, I c......

Prometheus Behind NAT: Scraping Without Open Ports

How a pull-based monitor reaches a machine with no public IP. The tunnel is a phone line, not a mailbox — and nothing exists until someone asks for it.

Prometheus pulls. It sends an HTTP GET to your machine and reads the response. Simple — until your machines sit in private subnets across three clouds, with no public IPs and security groups you’d ......

Load Leveling vs Load Balancing: The Fan-Out Pattern

Two queues, one event. Load balancing splits work across identical workers. Load leveling spreads work across time, so a slow consumer never blocks a fast one.

Most queue diagrams look the same: one producer, one queue, several workers. The point is to share work out. More workers, more throughput. Then one day you write a system where the same event goes......

Speculative Decoding and MTP: Why Guessing Is Free

A forward pass over five tokens costs about the same as over one. That gap is the entire speedup, and MTP is how the model drafts guesses to fill it.

I saw “MTP round-trip” on a checklist for a Megatron conversion pipeline and had no idea what it meant. Two acronyms, one hyphen, apparently important enough that someone had listed it as a thing t......

What special: true Actually Changes

It changes nothing when you tokenize. It changes everything when you detokenize, and only if skip_special_tokens is on. Measured against GLM 5.2's tokenizer.

I found the question in my own eval script, which is the embarrassing way to find things. 12text = tokenizer.decode(output[0, input_len:], skip_special_tokens=True).strip()text = re.sub(r"<......