Benchmarking Qwen 3.8 27B on RTX 5090 and beyond — VRAM capacity alone can't overcome severe software and inference engine bottlenecks
Detailed benchmarking of Alibaba's Qwen 3.8 27B open-weight AI model across multiple GPUs, including RTX 5090, RTX 4090, and RTX 3090. The review highlights severe bottlenecks in inference engines like llama.cpp and vLLM, showing that even with ample VRAM, prompt processing and token throughput can be poor. Dual RTX 5090 setups achieve up to 110 tokens per second with vLLM and MTP enabled, but at a cost of over $13,000. The article emphasizes that choosing the right model runner and hardware configuration is critical for local AI inference.
Related Tech News
AI CuratedMindfactory Data Shows 30% Jump In GPU Sales As Buyers Panic-Buy Ahead Of Another Looming Price Hike
Sales data from German retailer Mindfactory reportedly shows a 29.5% month-over-month increase in GPU sales in August, driven by buyers panic-buying ahead of possible price hikes. The Radeon RX 9070 XT remains the most popular GPU despite high prices. The article also cites Jon Peddie Research data showing a 10.4% quarterly increase in GPU shipments.
Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic
The Hugging Face blog post describes a research paper titled 'Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal.' It argues that current safety alignment treats harm as a topic-level property, but real deployments need to refuse only a subset of a topic (e.g., political manipulation) while answering benign prompts within the same topic. The paper formalizes this narrow-boundary safety setting and proposes a training method to achieve a sharp refusal step.
Mistral raises €3B as sovereign AI becomes big business
Mistral AI announced a €3 billion Series D round, the largest equity fundraising ever by a European tech company, led by Samsung Electronics. The funds will be used to scale compute capacity, build infrastructure, and expand international footprint. Mistral positions itself as a sovereign AI lab, aiming to address European concerns about dependence on US tech. The company has over 20 countries of operation and focuses on helping governments and corporations leverage AI.
AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome
Google DeepMind introduced AlphaGenome Atlas, a searchable platform containing predictions for 9 billion single-nucleotide variants in the human genome. It includes the AlphaGenome Variant Impact score and is available through a free portal, API, and Google Antigravity skill. The company says collaborators have already used it to identify variants in rare disease research.
