Skip to content
CB
  • Experience
  • Selected Work
  • Writing
  • About
  • Contact
  • Experience
  • Selected Work
  • Writing
  • About
  • Contact

Writing

How I spend my Friday nights. Build logs, benchmarks, and lessons from real systems.

  • Pooling a Tablet and a Workstation Over Thunderbolt to Run a 177B Model Wiring an RTX A6000 to a Ryzen AI Max+ 395 tablet over USB4 to run Qwen3.8-Flash-Next at full 262K context, and the control run that showed the cable made it slower.
    Sep 2026
  • TensorRT-LLM on an RTX A6000: Part 2 — Concurrency, TTFT, and the Benchmark Trap Pushing TensorRT-LLM under load on an RTX A6000: why synthetic prompt size misleads, why TTFT is a cleaner signal, and what it shows about Dynamo.
    Jun 2026
  • TensorRT-LLM on an RTX A6000: Part 1 — GPU Passthrough, KV Cache, and Single-GPU Serving Serving Qwen2.5-7B-Instruct with NVIDIA TensorRT-LLM on a single RTX A6000 — Docker GPU passthrough, KV cache, prefill/decode, Dynamo, and parallelism strategies.
    May 2026
CB

AI systems that are meant to be used, not just demoed.

  • Experience
  • Selected Work
  • Writing
  • About
  • Contact
  • Email
  • GitHub
  • LinkedIn
  • Resume
© 2026 Carl Broker · Duluth, Minnesota

Ask about Carl

Grounded in this site · runs on your device

Ask about Carl's projects, writing, tools, or background. Answers are built from this site's own pages — if something isn't covered here, this will say so rather than guess.

Powered by: Qwen3-0.6B · runs in your browser