Writing
How I spend my Friday nights. Build logs, benchmarks, and lessons from real systems.
- Pooling a Tablet and a Workstation Over Thunderbolt to Run a 177B Model Wiring an RTX A6000 to a Ryzen AI Max+ 395 tablet over USB4 to run Qwen3.8-Flash-Next at full 262K context, and the control run that showed the cable made it slower.Sep 2026
- TensorRT-LLM on an RTX A6000: Part 2 — Concurrency, TTFT, and the Benchmark Trap Pushing TensorRT-LLM under load on an RTX A6000: why synthetic prompt size misleads, why TTFT is a cleaner signal, and what it shows about Dynamo.Jun 2026
- TensorRT-LLM on an RTX A6000: Part 1 — GPU Passthrough, KV Cache, and Single-GPU Serving Serving Qwen2.5-7B-Instruct with NVIDIA TensorRT-LLM on a single RTX A6000 — Docker GPU passthrough, KV cache, prefill/decode, Dynamo, and parallelism strategies.May 2026