Kog Enables Processing of 3,000 Tokens Per Second with Standard GPUs
French startup Kog has announced a strategy to enhance AI inference speeds using standard data center GPUs. Kog has adopted an approach that involves designing software and model architecture together to maximize the utilization of NVIDIA and AMD GPUs. The company stated that it has integrated runtime, GPU kernels, and model architecture into an optimized pipeline to improve response speeds to requests from AI agents. According to a technology preview of the Kog inference engine released on May 28, it can process 3,000 output tokens per second per request using eight AMD MI300X GPUs, while eight NVIDIA H200 GPUs under the same conditions recorded 2,100 output tokens per second. Kog is primarily targeting AI coding agents and agent-based workflows. The Laneformer 2B model, released on Hugging Face on June 24, has 2.3 billion parameters and demonstrated performance of 45.1% on HumanEval+ and 51.6% on MBPP+. Kog reported in an August LinkedIn post that it achieved 2,857 tokens per second in a live demo. However, there is controversy regarding the interpretation of speed figures, and it has been pointed out that comparisons are difficult due to variations in model size and hardware. Kog's goal is to achieve low latency on existing GPU infrastructure.
-- Price
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.
You may also like

Tether AI Wallet Spending Limits, WDK Leaves Responsibility to Developers

ChatGPT, Claude, and Grok All Go Down: Why Is Everyone Suspecting Cloudflare?

Improvement in XRPL RPC Node Processing Capacity to 30,000 Messages per Second

Cyberattacks: ANSSI Can Now Impose Measures on Ministries with REACTIV
![[Interview] Stablecoins are Changing the Foreign Exchange Market — Barun Kumar, Co-founder of Hibachi.xyz](/public-static/26_2e1840f602.png?format=avif)
[Interview] Stablecoins are Changing the Foreign Exchange Market — Barun Kumar, Co-founder of Hibachi.xyz

US Court Seizes Some Virtual Assets Linked to North Korean IT Workers

European Commission Approves Purchase of Patriot Missiles for Ukraine Funded by €90 Billion Loan

Justin Sun Redeems 10,000 ETH from Lido Since August 26

Fake Friend Convinces Law Student to Transfer 1468 USDT

Ukrgasbank and Four Other Financial Institutions Fined by NBU: Largest Exceeds 30 Million UAH

Mortgage Loans for Self-Employed Workers: Requirements and Best Alternatives

Polish Prosecutors Charge Fifth Suspect in Zondacrypto Case

Solana’s “million-payments-a-second” AI system can leave sellers unpaid even after they deliver

Ukraine Appeals to OECD on Corporate Governance

Nepal Requests $2.5 Billion from UN Climate Fund After Devastating Floods

Arthur Hayes Releases FLOP Whitepaper, AI Inference Computing Power Commercialized on Blockchain

Hyperliquid Seeks U.S. Access with Permissive HIP-3

Crypto: Essential Tips for Developers to Protect Their Computers and Wallets

Ripple Swell adds former RBI governor Raghuram Rajan

Bloomberg: Crypto Exchanges Capture $800 Billion in Financial Market

Moonwell Card to Close on September 6, Users Must Withdraw Their Balances

Virtuals Announces Economic Infrastructure for AI Agents

Debate Over New ETF Regulations: Speed vs. Safety in the U.S. Cryptocurrency Industry

Washington State Proposes to Revoke Coinflip's License and Impose $1.0296 Million Fine

U.S. Government Asks Supreme Court to Activate New Mail-in Voting Rules

Zaprite Launches P2P Payment App in Bitcoin

Solana v1 upgrade bug can freeze network readers and disable fee limits

Flare Lowers Inflation to 3% and Implements Revenue Pool Stage

Cracking 1.33 Trillion Daily Tokens: B.AI Powers the “AI Grid” with Full-Stack Infrastructure to Fuel the Agentic Era











