Latest
Sep
29
AI agents shipped in 12 hour long Invide AI Agent Buildathon
On this Sunday, developers got together on invide Discord to turn their ideas into working projects across agent reliability, security,
4 min read
Sep
29
Claude Sonnet 5.5 is faster, but task cost depends on effort
Anthropic claims up to 30% lower cost per task. Artificial Analysis measured $7.60 per task at max effort, about 50% above Sonnet 5. Two request settings that worked with Sonnet 5 now return 400 errors.
3 min read
Sep
26
GLM-5.3 benchmarks, pricing and how it compares with Flash and rivals
GLM-5.3 keeps GLM-5.2's list price and matches Kimi K3's cost per Artificial Analysis index task. DeepSeek V4 Pro costs a third as much on the same index. Existing users must update requests that disable thinking.
8 min read
Sep
23
Opus 5.5, GPT-6 Sol, GPT-6 Luna and Grok 4.7 compared on cost
Opus 5.5's lower rates left max-effort task cost roughly level with Opus 5 in one independent test. Sol and Luna roughly halved their cost per task, while Grok 4.7 used more than twice the output tokens at unchanged rates.
4 min read
Sep
22
What are System One models? Jev and its open-source alternatives
TypeSafe's Jev returns typed answers and probabilities instead of text. Laya, Nimble, Kev and OpenThai now offer open versions of the same idea. This guide covers how they work, their limits and test results, and when a plain classifier is the simpler option.
7 min read
Sep
17
GLM-5.3-Flash costs up to 50x less than Opus. Your cost per finished task may not.
Given a billion tokens per vulnerability, GLM-5.3-Flash matched Claude Mythos Preview on ExploitBench at roughly 6% of the cost, while one production team found it fine for bounded work and not ready for their hardest tier.
10 min read
Sep
15
DeepSeek V4.1 Flash is a 552B model that bills like a smaller one
DeepSeek V4.1 Flash activates 8B parameters per input token and prices cached input from $0.003 per million. On Terminal-Bench 4.0, it scores 31.2 to Opus 5.0's 51.8.
6 min read
Sep
11
Shopify is leaving React Native after agents changed its cost model
Shopify rebuilt Shop in 12 weeks, backed by native specialists and strict parity checks. Maintaining two apps remains part of the bet.
3 min read
Sep
09
OpenAI says its AI solved Navier-Stokes. Codex users have another question
Navier-Stokes, a Millennium Prize problem in fluid dynamics, is still listed as unsolved. OpenAI denies accessing any specific user data but cannot rule out de-identified data from the researchers' Codex use improving its models.
4 min read
Sep
08
Jellyfin 12.0 is stable. Going back means restoring a backup
The release rewrites data on first boot, requires a full library scan, and leaves third-party plugins waiting for compatible builds from their maintainers.
3 min read