NVIDIA Qwen3.8-Flash-Next Achieves Over 16,000 Tokens/Second Throughput on GB300 NVL72
NVIDIA has announced that Alibaba's latest preview model, Qwen3.8-Flash-Next, is now supported on the NVIDIA GB300 NVL72 platform. The model has a total parameter scale of 176 billion, with approximately 6 billion parameters activated per token. It natively supports a context of 262,000 tokens and can be extended to 1 million tokens via YaRN, primarily targeting long-context agent applications such as intelligent programming, document processing, and tool invocation. NVIDIA stated that Qwen3.8-Flash-Next employs a mixed architecture of Gated DeltaNet (GDN) and Qwen Sparse Attention (QSA) to reduce computational and KV cache overhead in long-context scenarios. Testing shows that on the GB300 NVL72, the model achieves a single GPU throughput of over 16,000 tokens per second, with single-user throughput exceeding 200 tokens per second; it also supports inference frameworks such as SGLang, vLLM, and TensorRT-LLM.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Qualcomm Launches IMSDK 2.0 to Drive Edge AI Application Development

Apple Launches M6 Chip Mac mini Starting at $899

Thomson Reuters Develops Legal AI Model 'Thomson' with $40 Million Investment

EASY Residency Season 4 List Released, 9 Projects Have Interactive Opportunities

8 Years of 'Friendship': Internet Celebrity Di Shi Reveals He Was Defrauded of Millions by Sun Zeyu

Hash Global: After BTC, Who Will Take the Baton for the Next Bull Market?

Roman Storm's retrial postponed to April 2027

What Does the 'Niche Market' of Gold Mean?

Historically Unique: An Asset with 100% Winning Rate Over 4 Years

Potential Taxable Activities for Cryptocurrencies in 2025 Estimated at Approximately 73 Trillion Yen = Chainalysis

Credit Card Expenses in Dollars: The Best Way to Pay the Bill and Avoid Extra Charges

Anthropic AI Hardware Control Achieves 99.3% Success Rate in Tests

BTC Transfer from Kraken: 843 Bitcoins Leave the Exchange to Unknown Wallet

Euro Stablecoin on Ethereum: $579 Million, More Than Double Solana and Base

Alliance: AI Agents Are the Missing Piece of Web3

August 28 Cryptocurrency: Bitcoin May Surpass Gold's Market Cap, Predicts Binance Founder

DeFi Development Expands Holdings to 2.33 Million SOL with Resumption of Solana Purchases

"How I Almost Lost Half of My Bitcoins (and Kimi K3 Saved the Day)": Ledger User

Mirae Asset targets $109B digital asset business

SBI Invests 43 Billion Yen in Indonesia's Ajaib for Cryptocurrency Collaboration

U.S. Banking Regulators Exclude Reputation Risk, Reviving Debanking Discussions

$6.4 Billion in Bitcoin Options Expire Tomorrow—Here's What It Means

Stellar’s $3B RWA market faces a $2M DeFi gap

Why Ethereum Staking Will Change in 2026

Over 100 Companies Urge Strengthening of AI Cyber Defense

Mantle stablecoins and tokenized assets reach $880M

DefiLlama Ranks 128 Cryptos: Only One Receives AAA Rating

Villa Soldati: The CABA Neighborhood with the Highest Delinquency Rate

Mysterious $23 Million Bitcoin-Monero Swap: The Viral Video That Raises Questions





