July 2, 2026 · AI
Wiring codex to a Local Qwen — ollama, vLLM, and the developer role
Point OpenAI codex CLI at a local Qwen or gemma in five minutes
#LLM
한국어July 2, 2026 · AI
Point OpenAI codex CLI at a local Qwen or gemma in five minutes
June 25, 2026 · AI
MTP and diffusion inference on Gemma 4 and Qwen 3.6, fp8 on one H100
May 26, 2026 · AI
How harness design determines whether agents actually adapt.
May 22, 2026 · AI
Weights, prompts, and code as parameters at different layers of a learnable policy space
May 20, 2026 · AI
Benchmarking Qwen3.5-9B on Apple Silicon across MLX, llama.cpp, Ollama, omlx, and vLLM Metal — single-request throughput, prefill scaling, decode vs input length, and concurrency response
May 20, 2026 · AI
NIAH limits, the Lost in the Middle effect, alternative benchmarks, and measured recall across four reasoning-effort modes
May 20, 2026 · AI
NVIDIA's build.nvidia.com offers 100+ models on H100 infrastructure for free. Plug it directly into Claude Code, Cursor, or any OpenAI-compatible coding agent.