June 25, 2026 · AI
Speeding up LLM inference with MTP and diffusion
MTP and diffusion inference on Gemma 4 and Qwen 3.6, fp8 on one H100
#Benchmark
한국어June 25, 2026 · AI
MTP and diffusion inference on Gemma 4 and Qwen 3.6, fp8 on one H100
May 21, 2026 · AI
Eight graph engines measured across OLTP, memory, analytics, and differentiation queries
May 21, 2026 · DB
Why Arrow+GPU achieves 10x at SF=100 and Heap+GPU loses to CPU on wide tables
May 20, 2026 · AI
Five PDF-to-Markdown converters (markitdown, pdftotext, pymupdf, mineru, opendataloader-pdf) scored against a seven-criterion 100-point rubric
May 20, 2026 · AI
Benchmarking Qwen3.5-9B on Apple Silicon across MLX, llama.cpp, Ollama, omlx, and vLLM Metal — single-request throughput, prefill scaling, decode vs input length, and concurrency response
May 20, 2026 · AI
Implementing LightRAG's BaseGraphStorage on plain PostgreSQL with RCTE — why a 1-hop-dominant retrieval pattern fits flat SQL
May 20, 2026 · AI
NIAH limits, the Lost in the Middle effect, alternative benchmarks, and measured recall across four reasoning-effort modes
May 20, 2026 · AI
On a 1.14M-edge knowledge-graph workload, PostgreSQL RCTE beats Apache AGE by 290x — tracing the cypher() wrapper's 13ms cost and PG plan generation accumulation