x manual globalRelevance 90
AI Gateways Add New Models Immediately
Vercel AI Gateway added Grok Imagine Image 2.0 preview with direct AI CLI and playground access shortly after release.
Open sourcex manual globalRelevance 90
Cheaper Model Attempts Beat One Premium Run
Together AI reports that two DeepSeek V4 Flash attempts solved more DeepSWE tasks than one GPT-5.6 Luna attempt at roughly one-third of the cost.
Open sourceRedditRelevance 90
DeepSeek V4 Flash 0731 hits 82.7% on Terminal-Bench 2.1 in an independent public-harness run (445 trials)
Disclosure: I’m the author of Ante. DeepSeek recently reported an 82.7% score on Terminal-Bench 2.1 for DeepSeek V4 Flash 0731. Its evaluation used “DeepSeek Harness minimal mode,” which hasn’t been released yet. We wanted to see whether the reported result co...
Open sourceRedditRelevance 90
Updated benchmark: Deepseek V4 Flash on SlopCodeBench (local)
Howdy - I posted a benchmark here - https://www.reddit.com/r/LocalLLaMA/comments/1vbtiy7/deepseek_v4_flash_on_slopcodebench/ This was using the hosted API - since then I've been playing around with quants Here is the lastest benchmark - https://github.com/mich...
Open sourcehnRelevance 90
Qwen3.8 Max now ranked as the best overall model by agentic index
Comparison and analysis of AI models and API hosting providers. Independent benchmarks across key performance metrics including quality, price, output speed & latency. Artificial Analysis K Independent analysis of AI Understand the AI landscape to choose the b...
Open sourceRedditRelevance 90
Sorted last month's API spend by task type. The expensive part was all the boring work.
I exported last month's API usage and sorted by spend, expecting the hard stuff to be on top. It wasn't. Top of the list was classification, pulling fields out of documents, short summaries, and routing. Individually pennies. Collectively most of the bill, bec...
Open sourceDatabricksRelevance 90
Kimi K3 from Moonshot AI is now available on Databricks through Unity AI Gateway
A year ago, the best open-weight models trailed their proprietary counterparts by...
Open sourcex manual globalRelevance 90
We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE.
We analyzed DeepSeek-V4 Flash-0731 vs. GPT-5.6 Luna on software engineering tasks using DeepSWE. DeepSeek-V4 Flash-0731 delivers 80% of Luna’s performance at roughly 1/6 the cost. More insights in the thread!
Open sourcex manual globalRelevance 90
This graph from SiliconData is interesting. It's interesting to compare to cost per task, which actually shows the closed source models often to be even more expensive because they
This graph from SiliconData is interesting. It's interesting to compare to cost per task, which actually shows the closed source models often to be even more expensive because they're so token inefficient.
Open sourcehnRelevance 90
Cursor removed cost information from the usage page and CSV export
I just noticed Cursor Usage window switched from $$ to Token amount. I use this Usage window closely to keep tabs on my daily/active spending, not from the spending overall page. Today, the $$ amount is replaced by tok…
Open sourceRedditRelevance 90
Opus 5 dropped last week. We had it running in the business 10 minutes later, already pulling value, yes, as simple as that.
Half my feed was either panicking or acting like they'd built a new company overnight. Big tech CEOs make it like this, but a new model release doesn't fix a business that has nothing underneath it. If your AI advantage evaporates every time a new model ships,...
Open sourceHugging FaceRelevance 90
DeepSeek-V4-Flash-0731 Free Endpoint
Classic Mac-style guide to the free public DeepSeek-V4-Flash-0731 inference endpoint.
Open sourceHugging FaceRelevance 90
unsloth/DeepSeek-V4-Flash-0731-GGUF
Read our How to Run DeepSeek-V4-0731 Guide! Unsloth Dynamic 2.0 achieves superior accuracy & outperforms other leading quants. To run DeepSeek-V4-Flash-0731 in full precision lossless, run Q8 (UD-Q8 K XL), which is 162GB and only 7GB bigger than Q4 (UD-Q4 K XL...
Open source36KrRelevance 90
AI大厂,打起了“Token 奶茶大战”
过去两个月,AI公司突然开始集体“发券”。 北京时间7月31日,OpenAI宣布将GPT-5.6 Luna的API价格下调80%,Terra降价20%,Codex和ChatGPT Work中调用这两款模型消耗的额度也随之减少。 一时间,Token有了优惠券的既视感。它既可以是新品试喝券,也可以是系统故障后的补偿券。既能被装进会员套餐,也能做成“中杯、大杯、超大杯”,额度用完以后再续一杯。 这很像刚刚外卖界的“奶茶大战”:平台争相发券,看起来是在让用户占便宜,实际上争的是用户的消费习惯。 过去,AI公司的竞争主要发生...
Open sourcex manual globalRelevance 90
We are making the updated DeepSeek V4-Flash 0731 free in Cline.
We are making the updated DeepSeek V4-Flash 0731 free in Cline. This is the first flash model we've found performs at SOTA levels, and are excited for you to feel the new frontier. 1. npm i -g cline 2. Open /settings > Cline provider 3. Select deepseek-v4-flas...
Open sourcex manual globalRelevance 90
I burned 600M+ AI tokens last month and had no idea across which tools.
I burned 600M+ AI tokens last month and had no idea across which tools. So I built tkntracker — a local-first dashboard that tracks tokens from Claude Code, Codex, Cursor, Grok, Qwen, OpenCode & 20+ agents. No cloud. No API keys. No prompts leave your machine....
Open sourcehnRelevance 90
Launch HN: Tokenless (YC S26) – Automatic model switching to save money
Tokenless is a model router and drop-in replacement for the OpenAI and Anthropic APIs — the same quality at less than half the cost. Token less The router that cuts your inference bill in half . A drop-in replacement for your API calls — always routed to the r...
Open sourceRedditRelevance 90
Two weeks ago 62% of OpenRouter users spend happened on Anthropic models, that was reduced to 46% today.
Estimated spend per lab on OpenRouter Token consumption per lab on OpenRouter Token usage on OpenRouter continues to grow, although there is a steep decline in users dollar spend per lab in the last two weeks. This decline is mostly attributable to Anthropic m...
Open sourceHugging FaceRelevance 90
Prompt Routing
Route prompts with LFM2.5 Encoder on CPU
Open sourcehnRelevance 90
ARC-AGI Leaderboard
The ARC-AGI Leaderboard. ARC-AGI-3 Leaderboard ARC-AGI-1 ARC-AGI-2 ARC-AGI-3 Author: All Authors Model type: All Types Model: All Models Understanding the Leaderboard ARC-AGI has evolved from its first versions (ARC-AGI-1 and 2) which measured passive fluid in...
Open sourceRedditRelevance 90
Cheapest reliable API provider for open-source models (e.g., DeepSeek, GLM)
Hallo Everyone, i need to generate big amount of high quality data, for that i need some cheap API providers. who is the cheapest,most reliable provider you know of? submitted by /u/Zealousideal_Sort74 to r/LocalLLaMA [link] [comm...
Open sourceGitHub GrowthRelevance 90
lidge-jun/opencodex: +323 GitHub stars
Universal provider proxy for OpenAI Codex & Claude Code — use any LLM (Claude, Gemini, Grok, DeepSeek, Ollama…) with Codex CLI, App, SDK, and Claude Code
Open sourceAWSRelevance 90
Get started with OpenAI GPT-5.6 Sol, Terra, and Luna on Amazon Bedrock
OpenAI GPT-5.6 Sol, Terra, and Luna are now generally available on Amazon Bedrock. Learn how to select a model, run inference through the Responses API on the bedrock-mantle endpoint, reduce cost with prompt caching, connect the OpenAI Codex coding agent, and ...
Open sourcex manual globalRelevance 90
New research from Databricks shows that data agents can improve accuracy and lower cost at the same time.
New research from Databricks shows that data agents can improve accuracy and lower cost at the same time. We evaluated Genie Code against three leading general-purpose coding agents on 401 tasks distilled from real internal usage. Each agent used their own har...
Open sourcehnRelevance 90
Stripe in talks to buy OpenRouter for about $10B
Stripe is in discussions to acquire OpenRouter. Possibly for as high as $10 billion. Andrew Curran @AndrewCurran_ svg]:size-[1.125rem] text-subtext1 hover:bg-mix-current hover:bg-mix-amount-10 active:bg-mix-current active:bg-mix-amount-15 focus-visible:bg-mix-...
Open sourcehnRelevance 90
Show HN: Echo – Fable-level results at 1/3 the cost using open-weight models
I’ve been building Echo ( https://echo.tracerml.ai/ ), an experiment in making one AI system out of a pool of open-weight models rather than choosing a single model and using it for every task. It started with a simple experiment. I took a group...
Open sourceGitHub IssuesRelevance 90
can1357/oh-my-pi: Add built-in support for Requesty provider
### Description This issue tracks adding Requesty as a first-class provider in oh-my-pi's model catalog and AI registry, complete with its own authentication flow (omp login requesty) and proper model mappings. ### Use Case Users who already use Requesty as th...
Open sourceTechCrunchRelevance 90
Runway launches AI model router as generative media gets crowded
The Media Router is a tool that automatically selects the best image, video, or audio generation model for a request based on whether a developer prioritizes quality, speed or cost.
Open sourceDatabricksRelevance 90
Introducing AI spend controls with Unity AI Gateway
Today, we're announcing AI Spend Controls in Unity AI Gateway. This release extends...
Open source