Install
Build AI applications and inference infrastructure in chat. RunInfra selects open-source models, benchmarks GPUs, optimizes runtimes and kernels, and deploys production-ready AI endpoints.
- 2articles · 30d
- 4+ week agolatest article
- Aug 15, 2026earliest in window
- 0%with images
- 129avg words
- Computers & Electronics 2
- Science & Technology 2
- Software Dev. 1
- Software 1
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Qwen3.8 2.4T A95B | Model APIs
4+ week, 10+ hour ago (110+ words) Qwen3.8 2.4T A95B is an LLM listed in RunInfra Model APIs. RunInfra serves it as Inferact/Qwen3.8-2.4T-A95B-NVFP4 at $2.00 per 1M input tokens and $6.00 per 1M output tokens. Its context window is 262,144 tokens. The API provides OpenAI-compatible chat completions. USD, pay per token Hosted inference is available…...