Honest comparisons and stack breakdowns. No affiliate fluff.
NinjaTech bundled agents, inference, and reserved GPUs into one fixed annual fee priced in AI employees. Flat pricing is a ceiling, not a discount.
NaiveAI open-weighted a 309B MoE coding model under MIT with native 1M context and zero full-attention layers. What self-hostable long context actually costs.
Fireworks trained a model to stop over-thinking: same coding benchmarks as Kimi K3, 71% fewer reasoning tokens. Cost per task, not cost per token.
Autoheal's Evaluator scores agent runs against CI failures and incidents; its Healer opens PRs to change prompts and models. The eval loop is the product.
Gortex indexes your repo into a knowledge graph and serves it to coding agents over MCP, so they ask for three functions instead of grepping whole files.
Claude Code 2.1.283 adds deniedModels and availableModelsMatch managed settings, so a new model release can't enter your workflow until you approve it.
Stanford and NVIDIA released CLM-8B under Apache 2.0. It scores agent actions instead of generating them — 16.5ms decisions, and it caches the action side.
Anthropic is handing Pro and Max subscribers $100 and $250 in Claude Code credits that only work on cloud sessions, claimable to Oct 7, dead by Nov 4.
Docker published 11 Apache-2.0 SKILL.md skills that install into Claude Code, Codex, Cursor and Copilot. Vendor docs are becoming executable agent context.
Dataiku launched a cross-platform AI agent inventory on September 24. The premise is that most companies cannot list the agents they are already running.
Microsoft's new Copilot splits into Home, Code and Autopilot. The line that matters: background agents bill on usage, not on your seat.
OpenAI published MentalHealthBench — 1,215 conversations and 5,262 expert rubric criteria. Top score is 57.3%. The method is more useful to you than the leaderboard.
Databricks acquired Row Zero to put a governed spreadsheet in front of its Genie AI coworker. The lesson for small teams: your spreadsheet is the system of record.
Anthropic launched Claude Marketplace on September 23 with 2,000+ connectors and partner software you can buy with committed spend. What that changes for procurement.
Palo Alto's Continuous Frontier AI Defense runs Claude Mythos 5 and GPT-5.6-Cyber against your stack year-round. The multi-model finding is the part you can use.
A September 16 paper and an Apache 2.0 checkpoint run a 35B MoE from disk at ~15-20 tok/s on a 24GB Mac mini. What cheap local inference changes about your token bill.
Tencent's MIT-licensed BrowserSkill lets Claude Code, Cursor and Codex drive your real browser session. No test accounts — and no account boundary either.
Homebrew 7.0.0 adds a built-in vulnerability scanner, blocks home-directory access during builds, and closes eight advisories. What to run on your dev Macs today.
Anthropic folded Cowork into the main Claude app on September 16 and launched Claude Docs and Claude Slides in beta. The mode picker is gone — and so is the boundary.
GitHub shipped Copilot budget increase requests to GA on September 16. Usage-based AI billing now has an approval queue — and someone on your team has to own it.
GitHub Copilot's auto model selection now has Efficiency, Balance, and Intelligence tiers. You're billed for whatever it picks. Measure cost per finished task.
Socket Firewall in registry mode skipped upstream TLS verification by default. The fix landed June 29. The CVE landed September 12. Check your version.
A new benchmark runs eight frontier models against private production codebases. Top score is 38.8%. Why SWE-bench numbers don't predict your repo.
The huggingface_hub library tags every Hub request with the coding agent driving it. HF_HUB_DISABLE_TELEMETRY does not turn it off. Here is what leaks.
Salesforce introduced the Trusted Enterprise AI Harness on September 10, 2026. The unified experience rolls out in FY28. Here's what to build yourself now.
North Small Translate scores 83.60 on WMT26 against DeepL NextGen's 81.37, then ships under CC BY-NC. Read the license before you plan your localization.
SWE-2 scores 50.0% on FrontierCode 1.1 at 64% lower cost than Fable 5.1, post-trained on an open 2.8T Chinese base model. What that means for your coding agent budget.
VS Code 1.137 ships Automations in preview: recurring agent tasks that run hourly, daily, or weekly. Here's what that changes about cost and review.
NVIDIA's Sol-H3 stack renders 5 seconds of 1344x768 video with stereo audio in 1.653 seconds on 8x B300. AI video generation just crossed real time. Here is what changes.
Dependabot now reads private GitHub Packages with its built-in GITHUB_TOKEN. Delete the personal access token you wired into dependabot.yml years ago.
Stacklok's ToolHive is Apache 2.0 tooling that boots every MCP server in its own container with a permission profile and no local credentials. Here is the operator's read.
The Swiss Federal Chancellery will move ~3,000 staff to an open-source workplace by end-2027 for CHF 9M. Its proof-of-concept says what actually breaks.
GitHub Copilot content exclusions are now GA in the Copilot app and CLI. The docs still list gaps — agent modes, symlinks, and semantic leakage.
An AGPL-3.0 personal cloud launched September 5 with containerized apps, single sign-on, and zero telemetry. Self-hosting's real problem was never install.
The Virtual Disk Development Kit download pages went 404 in late August with no Broadcom statement. VDDK is how most migration and backup tools read VMware disks.
Fortune tracked GPT-6 Astra benchmark scores changing on OpenAI's live blog post after launch. Vendor pages are mutable — archive the numbers you decide on.
Google's Gemini Spark agent can edit, curate, and schedule recurring jobs across Google Photos. The scoping pattern is the part worth copying.
BugBase's Pentest Copilot Enterprise runs parallel agents across 100 vulnerability types without source code — and publishes its own benchmark results.
AMD revealed a Threadripper AI workstation with 576GB of HBM3E and named neither price nor availability. How to think about local AI inference hardware.
DoltLite swaps SQLite's B-tree for a prolly tree to give branches, merges and diffs on a SQL database — and an agent fleet did the compatibility work.
Appwrite 2.0 swaps proprietary abstractions for standard protocols — real PostgreSQL and MySQL, S3-compatible storage, and an OAuth 2.1 server. Portability as a product decision.
OpenClaw 2.0 landed with 16,000+ merged PRs, shared cloud sessions, and a rebuilt Control UI. The docs say the multiplayer controls are not tenant isolation. Read that before you roll it out.
Clipto raised $15M at a $250M valuation with 30M users and $15M ARR. The build lesson: local processing plus MCP beats uploading your archive to a vendor.
Claude Code 2.1.251 adds prompt-cache hit ratio and re-cached tokens to /cost. Cache misses, not prompt length, are what actually drive your agent spend.
VS Code 1.135 ships Rubber Duck, an experimental second-opinion pass from a complementary model, plus per-turn token accounting. Build the same loop in CI.
Anthropic is permanently raising Claude Code weekly limits 25% on September 14 — which is 17% less than the temporary boost you have today. Plan the delta.
Ten tech founders put $1M each behind a Linux distro's foundation. The interesting line isn't the money — it's that the foundation holds the trademarks.
Hugging Face's 25cm biped ships lidar, 15 motors, and the full reinforcement learning stack under Apache-2.0. Robotics prototyping is now an expense line, not a project.
Intel detailed Xeon 7 Diamond Rapids at Hot Chips 2026: 256 P-cores, PCIe 6.0, 1.6 TB/s memory. Shipping 2027 — which is a planning problem, not a purchase.
Particle launched Radar, a podcast search API and MCP server over 130,000 shows. The lesson for operators: agents can only use media somebody indexed first.
Nvidia's Groq 3 LPX rack benchmarked at ~3,400 tokens/sec, 4x the next-fastest endpoint. What token speed actually buys you in an agent workflow.
Z.ai shipped GLM-5.3-Flash under MIT — 320B total, 18B active, native vision, 1M context. The first cheap open-weight model aimed at document and screenshot work.
Google's Gemini Omni 1.1 Flash adds 360p drafts, 40-second scene extension, and 4K upscaling. How to build a video workflow that does not burn budget on takes.
Google shipped Gemini 3.5 Transcribe in public preview with 2.6% word error rate and 85+ languages. What accurate speech-to-text unlocks for a small operator.
Tencent led an $18M Series B into W4 Games, which sells enterprise support for MIT-licensed Godot. Buy the support, not the license — the pattern generalizes.
QueryStory surfaces the query behind every AI analytics answer and routes it for human review. The audit trail is the product — here's how to demand one.
OpenAI is putting the 5-hour rolling limit back on Codex and ChatGPT Work for Plus users. Your agent throughput is metered again — here's how to plan around it.
Apple's M5 Ultra Mac Studio tops out at 512GB unified memory and 1.2TB/s bandwidth, from $5,499. What that actually buys you for local model hosting.
AWS put GPT-5.6 Sol, Terra, and Luna inside Kiro and reported an 82% cost cut per completed Terminal-Bench task. The harness did that, not the model.
GitHub shipped a marketplace for MCP servers, plugins, skills, and canvases inside the Copilot app. Only the MCP server survives a change of vendor.
Ox Alpha is a free million-token coding model on OpenRouter from an anonymous provider that retains your prompts. Read that sentence twice before you route an agent at it.
TrueFoundry open-sourced TrueForge, an MIT-licensed agent harness. Benchmarked at $2.90 vs $11.80 against Claude Managed Agents. The loop is the part you should own.
Alibaba's Qwen-UI-Agent posts 82.1% on MobileWorld and 79.5% on OSWorld-Verified. The downloadable GUI agent is still the 2B and 8B predecessor.
Inherent's 27B Faraday agent beat Claude Opus 4.8 and GPT-5.5 at replicating research papers. What a small-model win means for your AI spend.
Firecrawl's Developer Index searches 70M+ repos, docs, issues, and merged PRs for coding agents. 0.63 recall@10 on DevDex. Why retrieval is your agent's real bottleneck.
Cloudways now hosts OpenClaw and Hermes as managed AI agents with one-click MCP access to your servers. The convenience is real. So is the blast radius.
Meta's new Mac app does system-wide dictation, screen context, and Google Workspace access. Read the training-data terms before your team installs it.
Cursor launched Origin, an agent-scale git forge, the same week GitHub threw errors for seven hours. How to add a remote without handing over your codebase.
Cerebras launched the CS-4 multi-wafer inference system claiming 30x GPU speed. The benchmark is real, the headline flops number needs an asterisk.
Microsoft's Intelligent Terminal 0.2 adds bring-your-own-model support, per-tab agent selection and WSL-native execution. A coding agent that never calls a cloud API.
Claude Code has been returning empty thinking blocks since mid-July while reasoning tokens still bill as output. The lesson is about vendor observability, not one bug.
ChatGPT Computer History records clicks, typing, and app switches on macOS for Pro, Business, and Enterprise. Off by default — here's the policy to set first.
DeepSeek open-sourced Harness v0.1 under MIT — a plugin-first agent runtime where models, tools, sandboxes and loops all swap in config. The harness is the product.
Cursor now boots Cloud Agents from pre-warmed environment snapshots instead of cloning and installing per run. The setup step was always billable — here's how to cut it.
Composio's new pricing lands August 15: the $29 tier drops from 200K tool calls to 50K, overage jumps to $4 per 1,000. Existing plans hold through December 31.
VS Code 1.133 opens the Agents window without GitHub sign-in and lets you swap model providers between turns. Your editor just got less locked in.
Microsoft is merging the consumer and Microsoft 365 Copilot apps and retiring Group Chats, Podcasts, Deep Research, and Copilot Labs. Export before the date.
Ryan Dahl's celld runs Cloudflare Workers and Durable Objects on your own S3 storage. The API is compatible, the license is Apache-2.0, the pricing claim is contested.
Automattic's Mesh CRM hit Android with an AI network query layer and a free 1,000-contact tier. What a personal CRM actually buys a small operator.
A researcher probed frontier models and found Claude Opus 5 answers like a January 2026 model despite a published May 2026 knowledge cutoff. Test your own.
A free 0-100 AI agent security score across six dimensions. Five of the six are ordinary supply-chain hygiene — which is the actual finding here.
GitHub Copilot for JetBrains now runs local Ollama models as a BYOK provider and remembers context across agent sessions. One is portability. One needs a policy.
Cross-session messaging lets Claude Code sessions hand off findings across terminals and machines. Here are the controls you should set before it matters.
GitHub shipped allowedMcpServers and deniedMcpServers for enterprise Copilot. Fail-closed by default. The pattern is worth copying even if you don't use Copilot.
Cloudflare shipped Kitesurf, a browser built for AI agents that uses 3-4x less CPU and up to 7x less memory than Chromium. Here's what it changes for your automation.
Nadella confirmed Microsoft is merging Copilot chat, GitHub Copilot, Cowork and Autopilots into one app. What consolidation does to your M365 bill and controls.
Atlassian's cloud revenue grew 31% while it guides Data Center down 17%. The self-hosted option ends March 28, 2029 — plan the migration on your calendar.
Reuters: the FCC is drafting an import ban on new Chinese optical transceiver models. The vendor with 27% share is Chinese. What that does to compute costs.
MacPaw is building on-device inference with Liquid AI and plans to hand it to Setapp developers. The credit meter is the part worth reading.
Microsoft Research's Orchard framework is MIT-licensed and public. Its own benchmarks show agents at 69.7% on SWE-bench and 59.6% on assistant tasks. Plan for that.
Genspark open-sourced GenOffice, an Apache 2.0 AI office suite for Mac and Windows. Read the routing: model calls go through Genspark's servers, not your key.
Cloudflare shipped a Billable Usage API for all self-serve accounts, FOCUS-aligned and live today. Why a queryable cloud bill beats a monthly PDF.
ShieldFont is a web font that feeds AI scrapers gibberish while humans see real text. Clever — and exactly wrong if you sell through AI search.
DeepSeek shipped V4-Flash-0731 on July 31 under an MIT license — 82.7 on Terminal Bench 2.1, $0.14 input and $0.28 output per million tokens. Route your cheap work here.
Cisco open-weighted two tiny models that find where a CVE lives in your codebase — Apache 2.0, 350M and 1B params, and your source never leaves the building.
PortSwigger's Burp AT puts AI agents inside Burp Suite Professional, with scope enforced outside the model. What agentic pentesting means for your web app.
Simile hit a $2B valuation selling AI stand-ins for consumers. Enterprises need simulation because they lost the customer. Small businesses haven't.
Pangram raised $9M from Menlo Ventures and shipped Pangram 4 plus an image detector. Every accuracy number is vendor-supplied. Here's how to test it on your own data.
Microsoft posted $90B in Q4 revenue and $115.9B in FY26 capex while pitching its own MAI models against OpenAI and Anthropic. Why the harness matters more than the model.
The FCC added humanoid robots, robot dogs, and power inverters to its Covered List on July 29, blocking new imports. What an import ban does to an automation roadmap.
A CVSS 8.7 n8n sandbox escape lets any workflow editor run OS commands on your host. Fixed in 2.31.5 and 2.32.1 — check your version today.
Ruff v0.16.0 raises the default rule set from 59 rules to 413 and formats Python in Markdown. Great defaults, and a very loud first CI run.
CodeMender hit public preview July 21: it finds a vulnerability, proves it's exploitable in your own sandbox, then writes the fix. Read the fine print.
GitHub made a 3-day Dependabot cooldown the default for version updates. Security updates still ship immediately. Why the delay is the point.
CVE-2026-29059 lets anyone read files off a Windmill server with no login. Patched in January, exploited in July. What self-hosting automation actually costs.
OpenAI's first hardware, a $230 keypad with status lights for AI coding agents, admits the real bottleneck: you can't tell what your agents are doing.
Anthropic's Claude Security plugin is in public beta for Claude Code — terminal vulnerability scanning with an adversarial verification pass. What it changes.
Runway launched a model router for generative media that picks by quality, speed, or cost. Useful — but routing logic is business logic you should own.
Claude Code Desktop can build, run, and tap through your iPhone app in a live simulator pane — no screen-recording permissions, but screenshots leave your Mac.
Abstract Security raised $25M for a streaming-first alternative to monolithic SIEM. The pattern — separate sources from destinations — applies far beyond security.
Katana shipped an MCP server and one-click AI replenishment for cloud inventory. Why a vendor-supported MCP endpoint matters more than the forecasting feature.
NVIDIA's SIGGRAPH 2026 news put MCP support into Adobe, Blender, Unreal and Houdini. Your content production pipeline just became scriptable by agents.
Infinity raised $15M to let an AI agent write inference kernels for any chip. The operator lesson: AI cost is a software layer, not just a hardware bill.
AWS launched CloudWatch coding agent insights for Claude Code, Codex, and Copilot. It's built on OpenTelemetry — which means you're not locked to AWS to get it.
VulnHunter is a free, Apache-2.0 agentic code security scanner from Capital One that hunts exploitable bugs instead of flagging patterns. Here's how to use it.
Epoch AI tested Pangram, GPTZero, and Originality.ai. Give a model a writing sample to imitate and detection collapses. Why you can't govern with AI detectors.
PyTorch 2.13 brings FlexAttention to Apple Silicon with a ~12x speedup on sparse patterns. Why that quietly lowers the cost floor for running your own models.
The Commerce Department moved the UAE to Country Group A:5, making advanced AI chip exports license-free overnight. What export policy means for where your AI runs.
Meta puts its custom Iris inference chip into production in September and is doubling compute to 14GW. What hyperscaler vertical integration means for your AI costs.
Claude Code desktop now has a sandboxed, isolated in-app browser the agent can read and click. The doc-lookup loop is gone — and browser-driving agents are now table stakes.
ZML released LLMD, a free inference server that runs open-weight LLMs across Nvidia, AMD, Google TPU, Apple, and Intel — breaking chip vendor lock-in.
SambaNova raised $1B at an $11B valuation for on-premises AI inference chips, with JPMorgan as a customer. Why owning the inference layer is now a bank-grade bet.
Squidbleed (CVE-2026-47729) is a Heartbleed-style memory leak in Squid proxy that survived 29 years. Here's who's exposed and what to do.
Baseten raised $1.5B at up to $13B for AI inference infrastructure. Why the boring serving layer — not the model — decides your cost and uptime.
Mistral OCR 4 adds bounding boxes, block types, and confidence scores at $2–4 per 1,000 pages — and runs in one container you control. What that unlocks for operators.
A single deterministic retrieval tool took biology-agent accuracy from 16.9% to 99.7%. The operator lesson: give agents real tools, not a bigger model.
Bhavin Turakhia is self-funding Neo, an AI-native, model-agnostic rival to Office and Workspace. The real lesson for operators: when to rebuild vs. bolt on.
Cursor shipped a native iOS app that fires off cloud coding agents to merge-ready PRs. What a small studio actually gets — and where the human checkpoint still goes.
X launched a hosted MCP server that lets Claude, Cursor, and Grok read the platform through your account — but not write. Why that read-only line is the right call for AI integrations.
We built Ceesvee, a CSV editor that handles 100MB+ files, on Tauri instead of Electron. Here's the honest comparison — including where Electron still wins.