French AI Startup Kog Bets on Software to Accelerate GPU Inference
AI inference acceleration competition is heating up. French startup Kog has opted not to develop its own chips but instead bets on optimizing underlying software to further unleash the inference capabilities of existing data center GPUs.
Targeting High Latency Scenarios in Enterprises
The company gained attention in May this year due to a technology preview. Kog showcased that using standard data center GPUs like AMD MI300X and Nvidia H200 can achieve extremely fast single-request decoding speeds. The company later stated that this demonstration generated about 200 clear business leads.
Kog is currently prioritizing professional workflows that require high response speeds, particularly in software engineering scenarios. According to the company, some heavy code generation users may wait for results for hours, and inference latency has become a significant issue affecting user experience and costs.
The company is also working with several design partners to advance application deployment. These clients allow users to generate games or applications through prompts, and improved inference speeds typically mean higher conversion efficiency and revenue potential.
Small Models Achieving 3000 TPS
Kog previously set a goal of "accelerating large model inference by 30 times," but current public demonstrations are still mainly based on small models. The demonstration used a Laneformer 2B model with about 2 billion parameters, achieving a single-request inference speed of 3000 tokens per second. This model is now open-sourced.
However, market demand has not remained focused on fine-tuning small models. Kog stated that potential clients are not prepared to focus on small models, so since the release of the demonstration, the company has shifted its main efforts toward accelerating the development of larger models to meet actual demand.
Validating Large Model Capabilities Around September
Kog believes that the potential of GPUs in inference decoding has not yet been fully realized. CEO Gaël Delalleau stated that as the memory bandwidth of the next generation of GPUs continues to improve, there is still significant room for optimization at the software level.
The cost of this approach is a longer R&D cycle. Kog stated that for each new GPU adaptation, the team often needs to invest several weeks or even months in hardware-level research. Given the current team size of 11, the number of chips the company can support in the short term remains limited.
Kog plans to gradually integrate this method into agent-based R&D processes in the future to support more chips and models. The more critical milestone now is to first prove that this approach can work with large models. Delalleau expects the company to achieve the first major model acceleration of about 10 times around September and use this to showcase client progress, driving subsequent Series A funding.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Crypto XRP: Evernorth Moves Closer to Wall Street and Nasdaq

Roman Storm's retrial postponed to April 2027

Credit Card Expenses in Dollars: The Best Way to Pay the Bill and Avoid Extra Charges

GoCaracal Attempts C2 Recovery Using Ethereum Storage Values

How to Quickly and Profitably Exchange USDT or USDC for Hryvnia in Ukraine

SimpleSwap Marks a Year Inside Exodus Wallet, Ships Five Partner Updates WithNo Integration Changes

DGrid AI: Reconstructing Trust Mechanisms in AI Infrastructure with PoQ and Verification Nodes

Frostsnap Releases Version 0.4.0, Fixes Two Security Issues

Connecticut sues Kalshi over unlicensed sports event contracts

FDA approves Rasonque, the first oral therapy targeting the RAS pathway in pancreatic cancer

Rent Tron Energy to Reduce USDT Transaction Fees: TronBid Expands Trading Market

How to Have Internet Abroad Without Spending Hundreds of Dollars: The Difference Between eSIM and Roaming

Bitcoin Core reduces optional txindex size by 40 GB

Linux 7.3 Boosts FUSE Performance with Buffer Groups and Zero-Copy

Linux 7.3 Adds Support for Non-Fatal PCI Errors and New Intel Chips

Unstoppable Domains drops ICANN plans, offers refunds

Injective Expands AI Payment Experiment by Joining X402 Foundation

Former Saitama CEO Manpreet Kohli Faces US Extradition Over $20 Million Fraud Charges

Arthur Hayes: The Eligibility for FLOP Token Airdrop Will Depend on Users' Testnet Activity

Saitama Founder’s Extradition Case Rejected, May Be Extradited to the U.S. for Trial

Zelensky Awards Elon Musk the Order of Freedom

BNB Chain x402 Payments Exceed 1.4 Million Transactions

Tornado Cash Co-founder Roman Storm's Retrial Postponed to April 2027, Court to Examine 'Is Code a Crime?'

SMEs: The Government Launched the RIMI Six Months After Its Approval in Congress

Vietcombank Warns of AI-Driven Fraud

Chile Requests the Extradition of Two Venezuelans Who Laundered Nearly USD 19 Million for the Tren de Aragua

Firefox Discovers 40 Malicious Cryptocurrency Wallet Extensions Stealing Mnemonics

Interview: Aptos CEO on AI Agents, Stablecoins, and Machine Commerce

Former Deputy Head of the President's Office Mudra Sent to Pretrial Detention with a Bail of 20 Million UAH








