鈴木商店
鈴木商店 SUZUKI SHOTEN
Home Projects Series Blog Analysis Archive 🇯🇵 JP Contact
🇯🇵 JP
Home Projects Series Blog Analysis Archive
Contact on LinkedIn
Home / Blog / GB10
GB10

2 articles tagged with "GB10"

Part 4. TANREN — an Evolved Scheduler Patched into vLLM Halves Tail Latency Under Heavy Load Series

Part 4. TANREN — an Evolved Scheduler Patched into vLLM Halves Tail Latency Under Heavy Load

The scheduler — the heart of LLM serving — evolved by TANREN and A/B tested inside a real vLLM. Across 16 real-trace configurations it beats the default FCFS 15 times (p=2.6e-4), cuts mean/p99 TTFT by 35-49% under heavy load, and does no harm when idle. Includes the honest record of going 0/6 on unseen data and six rounds of diagnosis before the win.

2026-07-25 Read →
Part 3. TANREN — Patching the Evolved Cache Policy into a Real vLLM and Measuring It Series

Part 3. TANREN — Patching the Evolved Cache Policy into a Real vLLM and Measuring It

The cache policy that won in simulation is patched into a live vLLM and measured on real hardware. The simulated win vanished at first; after fixing where and what to measure, tail latency (p99 TTFT) came out 13-16% lower under memory pressure. The prototype's costs and the conditions where it does not help are reported as measured.

2026-07-25 Read →

All Tags

AI (62) AI Agent (7) AI Agents (3) AI Cost Optimization (2) AI Infrastructure (6) AI Security (1) AI Strategy (1) AWS (3) Advertising (2) Agent (3) Alibaba (2) Amazon (3) Android (2) Anthropic (5) Arm (1) Atari (5) Bluetooth (1) Business (17) ByteDance (1) CPU (1) CXL (1) Caching (2) China (3) Claude (3) Cloud (1) Cloudflare (1) Code Generation (1) Coinbase (1) Computer Vision (8) Cost Management (1) Cost Optimization (1) Cost per Task (1) Crypto (1) Cryptocurrency (1) Cybersecurity (2) DRAM (1) Data Center (5) Datadog (1) Earnings (1) Edge AI (1) Edge Computing (1) Education (3) Elon Musk (1) Energy (2) Enterprise AI (1) Enterprise Software (1) Ethics (1) Evolutionary Search (8) FDE (1) Finance (2) Fine-tuning (6) Fintech (1) Firestore (1) FunSearch (6) Future of Work (1) GAN (7) GB10 (2) GPU (4) GTC (1) Game Dev (3) Gemini (4) Gemma3 (1) Generative AI (3) Geopolitics (1) Git (1) GitHub (1) Google (2) Google Cloud (2) Google Colab (1) Google Maps (2) HBM (1) Hetzner (1) Hyperscaler (1) IPO (4) Incident (1) Inference (1) Infrastructure (16) Intel (1) Investing (2) Investment (12) Java (2) Jevons Paradox (1) KOI (2) LLM (21) LLM Comparison (1) Leveraged Loans (1) LoRA (2) M&A (1) Machine Learning (14) Memory (1) Meta (4) Micron (1) Microsoft (3) Mobile (2) Model Routing (1) Multi-Agent (1) MySQL (1) Mythos (2) NLP (6) NVIDIA (5) Neocloud (1) Next.js (5) Nishiki (1) Ollama (1) Open Source (1) OpenAI (9) OpenClaw (5) OpenRouter (1) PHP (5) Palantir (1) Payments (1) Phaser 3 (3) Prediction Market (1) Python (24) Real-Time (2) Reinforcement Learning (5) SK hynix (1) SaaS (5) Samsung (1) Scheduling (1) Security (2) Semiconductor (3) Semiconductors (2) Sensors (2) SoftBank (2) Space (2) SpaceX (4) Stablecoin (2) Starlink (1) Stock Prediction (5) Strategy (1) Streaming (1) Stripe (1) System Design (1) TANREN (8) TTS (1) Tencent (2) Tesla (2) Text-to-SQL (1) Tokenmaxxing (2) Translation (2) Twitter API (1) TypeScript (3) Vertex AI (1) Virtual Try-On (3) White-Collar (1) Windows (1) heartbeat (1) iOS (1) vLLM (3) xAI (2)
鈴木商店

鈴木商店

SUZUKI SHOTEN

MarketQuest Nishiki
© 2026 鈴木商店 / SUZUKI SHOTEN All rights reserved.
Search articles...