标签
#inference
19 条相关内容
Predictive Speculative KV Replication for Bursty LLM Inference
🔥Hacker News2026-07-31
ChatGPT Optimizes Its Agent Loop: Harness, API, and Inference
🔥Hacker News2026-07-30
Kimi K3 Now Available via Telnyx Inference API
🔥Hacker News2026-07-27
Ask HN: HotPin – lossless 120B MoE inference on 24GB RAM (CPU, 50 loc)
🔥Hacker News2026-07-25
Hetzner is working on LLM Inference
🔥Hacker News2026-07-24
AMD and Cerebras Launch AI Inference Solution
🔥Hacker News2026-07-24
Show HN: Flow Matching model inference in C
🔥Hacker News2026-07-23
kvcache-ai/ktransformers
A Flexible Framework for Experiencing Heterogeneous LLM Inference/Fine-tune Optimizations
GitHub2026-07-21
Homomorphically encrypted CIFAR-10 inference in 200ms
🔥Hacker News2026-07-17
Accelerating Block Low-Rank Foundation Model Inference on MemoryConstrained GPUs
🔥Hacker News2026-07-16
Performant C/CUDA inference engine for Qwen 3.6 35B on RTX 5090 / Blackwell
🔥Hacker News2026-07-13
Litert.js, Google's High Performance Web AI Inference
🔥Hacker News2026-07-12
Show HN: Reame – a CPU inference server that gets faster as it runs
🔥Hacker News2026-07-11
Inference Optimization for MiMo v2.5: Pushing Hybrid SWA Efficiency to the Limit
🔥Hacker News2026-07-07
AI inference is obviously profitable
🔥Hacker News2026-07-03