Inference
Free
vLLM
Adoption
GitHub forks
19.3K▲ 1.6%
GitHub stars
86K▲ 0.7%
PyPI / week
1.2M▲ 0.7%
Objective usage signals from public registries (Hugging Face, GitHub, npm, PyPI, Wikipedia), captured every 6 hours. Δ compares to ~7 days ago.
Why is this hot
6 mentions / 7d across 2 sources (diversity ×1.15)last 48h: 0 vs baseline 1.5/day
- New OllamaMQ v0.3.0hn
- Native vLLM and ROCm 7.15 for RX 6000 (RDNA2) on Windows 11 – 26 Tflops FP16hn
- AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Earxiv
- Write Once, Run Everywhere: The Axon DSL for Shape-Safe and Framework-Agnostic Larxiv
- ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedarxiv
Top mentions by weighted contribution (source authority × engagement × recency).
Description
High-throughput LLM serving engine.
Recent mentions in news (15)
- AsmEvo: Agentic Assembly-Level Optimization of AMD GPU Kernels with Functional Equivalence Verificationarxiv-ai · 3d
- Write Once, Run Everywhere: The Axon DSL for Shape-Safe and Framework-Agnostic LLM Architecturesarxiv-ai · 4d
- Flama: a Python framework for development and deployment of production-ready APIs, machine learning, and LLM servicesarxiv-ai · 5d
- New OllamaMQ v0.3.0hn-ai · 5d
- Native vLLM and ROCm 7.15 for RX 6000 (RDNA2) on Windows 11 – 26 Tflops FP16hn-ai · 6d
- ArguLens: An Open-Source System for Automated Essay Scoring and Label-Aware Feedback Generationarxiv-ai · 6d
- Pallas: A Proactive KV Cache Migration Framework for LLM Inference in AI-RANarxiv-ai · 7d
- Beyond Binary Priorities: Multi-Tier SLA Scheduling for Large Language Model Servingarxiv-ai · 7d
- Engineering Signals of Human-AI Collaboration in the Agentic Coding Era: A Longitudinal Analysis of 33,228 Pull Requests from vLLM and SGLang with Implications for Biomedical AI Agents and Bioinformatics Pipeline Developmenarxiv-ai · 10d
- The Integer Alibi: Localizing Cross-Kernel Divergence in INT8-Quantized LLM Inferencearxiv-ai · 10d
- Show HN: NanoRL – RL training for LLMs in ~1,800 lineshn-ai · 11d
- vToken: Token-Level Virtualization for Reclaimable KV Cachesarxiv-ai · 11d
- Trie Automata for Constrained Decoding over Large Finite Setsarxiv-ai · 11d
- Ask HN: Are there any production LLM pipeline setups to learn from?hn-ai · 11d
- Scheduling Mixed RL Rollouts Beyond Prefix Localityarxiv-ai · 12d
Related Briefs & Stories (6)
8.2
1 items • 27dinbox
7.3
1 items • 1moinbox
7.1
1 items • 1moinbox
7.1
1 items • 2moinbox
7.0
1 items • 28dinbox