tomshardware.com··review

Benchmarking Qwen 3.8 27B on RTX 5090 and beyond — VRAM capacity alone can't overcome severe software and inference engine bottlenecks

AI Quality: 87/100Freshness: 98/100
Benchmarking Qwen 3.8 27B on RTX 5090 and beyond — VRAM capacity alone can't overcome severe software and inference engine bottlenecks
Key Takeaway

Detailed benchmarking of Alibaba's Qwen 3.8 27B open-weight AI model across multiple GPUs, including RTX 5090, RTX 4090, and RTX 3090. The review highlights severe bottlenecks in inference engines like llama.cpp and vLLM, showing that even with ample VRAM, prompt processing and token throughput can be poor. Dual RTX 5090 setups achieve up to 110 tokens per second with vLLM and MTP enabled, but at a cost of over $13,000. The article emphasizes that choosing the right model runner and hardware configuration is critical for local AI inference.

AI Summary & Analysis
Tom's Hardware benchmarks Qwen 3.8 27B on RTX 5090 and other GPUs, finding that VRAM alone is insufficient due to software and inference engine bottlenecks.
Original Source Coverage
tomshardware.com
Read original story at tomshardware.com
Related Topics:
#ai#benchmarks#gpu#llm#models#nvidia#qwen

Related Tech News

AI Curated
wccftech.com09/08

Mindfactory Data Shows 30% Jump In GPU Sales As Buyers Panic-Buy Ahead Of Another Looming Price Hike

Sales data from German retailer Mindfactory reportedly shows a 29.5% month-over-month increase in GPU sales in August, driven by buyers panic-buying ahead of possible price hikes. The Radeon RX 9070 XT remains the most popular GPU despite high prices. The article also cites Jon Peddie Research data showing a 10.4% quarterly increase in GPU shipments.

newsRead summary →
huggingface.co09/08

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

The Hugging Face blog post describes a research paper titled 'Safety for Whom? Boundary-Aware Self-Distillation for Controlled LLM Safety Refusal.' It argues that current safety alignment treats harm as a topic-level property, but real deployments need to refuse only a subset of a topic (e.g., political manipulation) while answering benign prompts within the same topic. The paper formalizes this narrow-boundary safety setting and proposes a training method to achieve a sharp refusal step.

newsRead summary →
techcrunch.com09/08

Mistral raises €3B as sovereign AI becomes big business

Mistral AI announced a €3 billion Series D round, the largest equity fundraising ever by a European tech company, led by Samsung Electronics. The funds will be used to scale compute capacity, build infrastructure, and expand international footprint. Mistral positions itself as a sovereign AI lab, aiming to address European concerns about dependence on US tech. The company has over 20 countries of operation and focuses on helping governments and corporations leverage AI.

newsRead summary →
deepmind.google09/08

AlphaGenome Atlas: A predictive map of every possible DNA letter change in the human genome

Google DeepMind introduced AlphaGenome Atlas, a searchable platform containing predictions for 9 billion single-nucleotide variants in the human genome. It includes the AlphaGenome Variant Impact score and is available through a free portal, API, and Google Antigravity skill. The company says collaborators have already used it to identify variants in rare disease research.

newsRead summary →