拳禅一体 Fist and Zen as One
色即是空 Form is Emptiness
自他共栄 Mutual Prosperity
武魂洋才 Samurai Spirit, AI Frontier

Portfolio

Blog

Insights & Analysis

Part 1. What Makes the New AI Model "Jev" Special? Reproduced with NerveReflex and an Open Model Series
AIJevNerveReflexHidden StateGemmavLLMInference CostCalibration

Part 1. What Makes the New AI Model "Jev" Special? Reproduced with NerveReflex and an Open Model

Jev is a new decision-focused AI model announced by ChatGPT co-inventor Diogo Almeida. This post separates the announcement from social media claims, explains why the approach is fast and cheap, and shows how I reproduced the same kind of setup with NerveReflex ideas and the open model Gemma 4 26B.

Read more →
Why AI Keeps Slipping Out of Control: What the OpenAI and Anthropic Incidents Show About the Difficulty of Containment Insight
AI SafetyAnthropicOpenAIClaude CodeRecursive Self-ImprovementAI Agents

Why AI Keeps Slipping Out of Control: What the OpenAI and Anthropic Incidents Show About the Difficulty of Containment

Running Claude Code autonomously for long sessions and chaining multiple sessions together made AI safety in the age of Recursive Self-Improvement feel concrete to me. Using the recent incidents at Anthropic and OpenAI, I look at why monitoring the chain of thought or sandboxing alone is not enough, and what a realistic layered defense looks like.

Read more →
OpenAI's Next Model "Astra" — Thinking Deeper While Cutting Inference Cost, and the Security Concerns That Come With It Insight
AIOpenAILatent ReasoningRecurrent DepthNerveReflexAI SafetyInference Cost

OpenAI's Next Model "Astra" — Thinking Deeper While Cutting Inference Cost, and the Security Concerns That Come With It

Recurrent Depth is reported to be used in OpenAI's next model, Astra. I look at what it means to keep reasoning inside the model's internal representation instead of writing it out as words — for inference cost, for AI safety monitoring, and for NerveReflex, the architecture I have been building.

Read more →
Preventing Enterprise AI Agent Failures — the "20% AI, 80% Deterministic" Design I Learned in the Field as an FDE Insight
AIAI AgentsEnterprise AIFDELLM

Preventing Enterprise AI Agent Failures — the "20% AI, 80% Deterministic" Design I Learned in the Field as an FDE

Lessons from building HR and finance AI agents in the field as a forward deployed engineer: the 20% AI / 80% deterministic design, the limits of general-purpose assistants, and why coexistence beats full automation.

Read more →
Part 7. TANREN — Atari Bowling, the Lost Campaign That Built the Entry Gate: Weighing a Fight Before Running It Series
TANRENEvolutionary SearchFunSearchAtariReinforcement LearningLLMPython

Part 7. TANREN — Atari Bowling, the Lost Campaign That Built the Entry Gate: Weighing a Fight Before Running It

On Atari Bowling, TANREN beat R2D2, the human benchmark, and DQN, then hit a ceiling at 74% of the strongest RL. This is the record of establishing — by measurement — that the ceiling was structural to the class of readable reactive code, and of how that loss produced the entry gate that now guards every experiment.

Read more →
Part 5. TANREN — Discovering Breakout's Tunnel Strategy on Its Own, and Measuring Exactly Why the Rest Is Out of Reach Series
TANRENEvolutionary SearchFunSearchAtariReinforcement LearningLLMPython

Part 5. TANREN — Discovering Breakout's Tunnel Strategy on Its Own, and Measuring Exactly Why the Rest Is Out of Reach

In Breakout, evolution discovered the classic tunnel strategy without being taught it and reached about 6.1x the human benchmark. Just as important, the reason it could not reach the strongest RL agents was pinned down by measurement, not speculation — drawing the boundary of what a few dozen readable lines of code can do.

Read more →
Part 2. TANREN — 11 Wins to 1 Against Twenty Years of Cache Classics, on the Future of a Production Trace Series
TANRENEvolutionary SearchFunSearchCachingLLMPythonvLLM

Part 2. TANREN — 11 Wins to 1 Against Twenty Years of Cache Classics, on the Future of a Production Trace

Cache eviction — deciding what to drop when a cache is full — has been studied for more than twenty years. On that ground, a 53-line evolved Python function beats ARC, LIRS, and W-TinyLFU on future segments of Twitter's production traces. This post covers how the first attempt lost by overfitting to one time period, the two-window training that fixed it, and the public code anyone can rerun.

Read more →
Part 8. TANREN — The Number I Almost Claimed and Then Withdrew: Atari Freeway and Measurement Honesty Series
TANRENEvolutionary SearchFunSearchAtariReinforcement LearningLLMPython

Part 8. TANREN — The Number I Almost Claimed and Then Withdrew: Atari Freeway and Measurement Honesty

A score of 31.9 looked like a DQN-beating result — until re-measuring under the published side's exact protocol dropped it to 28.6. This is the record of withdrawing the number, repairing the verifier, and re-running the evolution under the correct protocol to reach 31.2: past the human benchmark and DQN, at 96% of Agent57. An article about evaluation honesty.

Read more →
Part 6. TANREN — On Atari Skiing, Deep RL's Hardest Slope, Readable Python Beats Every Published RL Agent Series
TANRENEvolutionary SearchFunSearchAtariReinforcement LearningLLMPython

Part 6. TANREN — On Atari Skiing, Deep RL's Hardest Slope, Readable Python Beats Every Published RL Agent

On Atari Skiing — where Agent57 needed roughly 78 billion training frames to surpass the human benchmark — about 130 lines of readable Python scored −3310.7 in a held-out evaluation, beating every published RL agent and the human benchmark. Includes the entry gate that estimated the odds before running, the techniques evolution added on its own, and the cost contrast: one day and a few dollars.

Read more →

Products

Publicly available services built with AI and modern web technologies.