majlis@daily — 16 Sep 2026 — Daily AI News — AI Majlis Daily
majlis@daily:~$ cd ~ / articles / 2026-09-16

Daily AI News — 16 Sep 2026

# Wednesday, 16 September 2026 · the 10 stories that matter, ranked
# jump to a story
01OpenAI, Anthropic & Google DeepMind hold weeks of talks on an AI safety standards body [Policy]The three biggest labs are quietly coordinating on frontier-model safety testing.02Factory raises $200M at a $5B valuation for enterprise coding agents [Funding]An AI coding-agent startup just tripled its valuation in five months.03OpenAI's custom “Jalapeño” chip beats Nvidia Blackwell on efficiency in first benchmarks [Hardware]OpenAI's first in-house inference chip posts industry-leading numbers.04Enterprise AI-agent security funding hits $435M in five months [Funding]VCs are pouring money into making agents safe enough to actually ship.05Clay doubles to a $7.1B valuation as sales-automation agents heat up [Funding]Sales-and-marketing automation is the other agent gold rush.06Google DeepMind's new chief Koray Kavukcuoglu inherits the race to catch OpenAI & Anthropic [Industry]A leadership change at DeepMind, with coding as the battleground.07Musk proposes AI labs peer-review each other's models [Policy]A rival-tests-rival idea enters the safety debate.08Euno raises $23M to give AI agents enterprise “context” [Funding]The unglamorous layer agents need: knowing what your data means.09Skild AI unveils S1, a robot that learns 10-minute tasks from a single video [Robotics]A robotics foundation model that generalizes from one demonstration.10Apple debuts its first 2nm M6 chip, pushing on-device AI [Hardware]Local AI gets a meaningful hardware bump.
#01 · Policy

OpenAI, Anthropic & Google DeepMind hold weeks of talks on an AI safety standards body

# The three biggest labs are quietly coordinating on frontier-model safety testing.
> the_news

OpenAI global-policy chief Chris Lehane confirmed the company has spent several weeks in talks with Anthropic and Google DeepMind on coordinating AI safety, following Dario Amodei's essay urging labs to slow frontier development. The discussions centre on a proposed US-led “Frontier AI Standards Body” — modelled on FINRA — that would test the most capable models before release, focused on cyber-offense, bio-misuse and deception. No regulator, testing requirement or deployment rule has actually been created yet; the pre-release access window (up to 30 days) would start voluntary. Lehane said the labs do not believe they need an antitrust waiver to collaborate.

> why_it_matters

This is the first time direct rivals have openly coordinated on safety governance — a signal that self-regulation, not legislation, is the near-term path. For anyone building on these models, it hints at future pre-release review cycles that could affect launch timing and available capabilities.

> real_life_use_case

If you ship products on frontier APIs, start tracking each lab's safety-eval posture now — model release cadence and feature gating may increasingly hinge on these voluntary reviews.

#02 · Funding

Factory raises $200M at a $5B valuation for enterprise coding agents

# An AI coding-agent startup just tripled its valuation in five months.
> the_news

Factory, which builds AI agents for enterprise engineering teams, said it raised $200M in a round that more than tripled its valuation to $5B (up from $1.5B in April 2026). Reuters reported the round was co-led by Blackstone, Khosla Ventures, Sequoia, Insight Partners, Evantic Capital and Sound Ventures. Founded in 2023 by Matan Grinberg and Eno Reyes, Factory competes directly with Cognition and Cursor, covering software design, testing and support under its “software factory” model — a system that turns bug reports, internal comms and customer feedback into shipped code.

> why_it_matters

AI-assisted coding is now the clearest killer use-case in generative AI, and investors are repricing agent startups at extraordinary speed. A $5B valuation on a two-year-old company shows how much capital believes engineering work is being automated.

> real_life_use_case

Engineering leaders should pilot an agentic coding tool on a real backlog this quarter — not just autocomplete, but end-to-end “ticket → PR” flows — to benchmark how much of routine dev work can be delegated.

#03 · Hardware

OpenAI's custom “Jalapeño” chip beats Nvidia Blackwell on efficiency in first benchmarks

# OpenAI's first in-house inference chip posts industry-leading numbers.
> the_news

At Hot Chips, OpenAI presented the first independent benchmarks for Jalapeño, its custom inference ASIC co-designed with Broadcom. On SemiAnalysis's InferenceX benchmark, it delivered 1.5–1.9× more work per watt than an Nvidia Blackwell system and 1.7–3.6× lower end-to-end latency, running on ~550W sustained vs Blackwell's far higher draw. SemiAnalysis called it “industry-leading,” beating every Nvidia, AMD and Google chip they tested — rare for a first-gen part. Design-to-tape-out took roughly 16 months; OpenAI used its own models to help accelerate the hardware engineering. Small-scale rollout is expected late 2026, broader scaling in 2027.

> why_it_matters

Power — not silicon — is the binding constraint on AI buildouts. A chip doing equal work at half the watts lets the same data-center substation serve twice the paying customers, and loosens the industry's dependence on Nvidia.

> real_life_use_case

If your product's margins are eaten by inference cost, watch for custom-silicon capacity opening up in 2027 — it could meaningfully cut per-token pricing for latency-sensitive agent workloads.

#04 · Funding

Enterprise AI-agent security funding hits $435M in five months

# VCs are pouring money into making agents safe enough to actually ship.
> the_news

Between April and September 2026, investors put $435M into 12 financings for enterprise AI-agent security and governance startups — nine focused solely on making agents safe to run inside businesses. Zenity raised a $125M Series C (Norwest-led); Alice raised $140M (Apax) and is nearing $100M ARR with eight of the ten leading labs as customers; AIR came out of stealth with $50M for a runtime “firewall” that vets the skills, plugins and MCP servers agents use. The driver: IDC/Lenovo research found 88% of enterprises with agent initiatives never reach production.

> why_it_matters

The gap between agent hype and production is a governance problem, and this is the infrastructure being built to close it. It confirms security — not model quality — is now the bottleneck to enterprise agent adoption.

> real_life_use_case

Before deploying agents on real systems, add a runtime guardrail layer (action monitoring, skill/plugin vetting) — it's fast becoming the default checklist item for moving from pilot to production.

#05 · Funding

Clay doubles to a $7.1B valuation as sales-automation agents heat up

# Sales-and-marketing automation is the other agent gold rush.
> the_news

Clay, whose software automates sales and marketing tasks, raised $115M at a $7.1B valuation — more than double its worth a year earlier. The round reflects intensifying investor appetite for application-layer AI agents that do concrete revenue work: enriching leads, researching accounts and automating outbound. It lands alongside the broader $435M wave into agent security infrastructure, showing the market building both the agents and the guardrails to govern them in parallel.

> why_it_matters

GTM is one of the fastest-moving frontiers for practical AI — agents that touch pipeline and revenue get funded fastest because ROI is measurable. For go-to-market teams, this is where AI is changing daily workflows now, not eventually.

> real_life_use_case

Sales and growth teams can wire a lead-enrichment/research agent into their CRM this quarter to auto-build account briefs before every call — a low-risk, high-leverage first agent deployment.

> source

Edgen ↗

#06 · Industry

Google DeepMind's new chief Koray Kavukcuoglu inherits the race to catch OpenAI & Anthropic

# A leadership change at DeepMind, with coding as the battleground.
> the_news

Koray Kavukcuoglu — previously DeepMind's CTO and Google's chief AI architect — is becoming SVP and head of DeepMind, reporting directly to Sundar Pichai, as Demis Hassabis moves to chair. He'll oversee Gemini development, frontier research and the Gemini app. Analysts told CNBC Google fell behind because it focused on Search and multimodal and missed the first killer use-case — coding — where OpenAI and Anthropic are “miles ahead.” Google hasn't shipped a frontier-leading model since early 2026 (Gemini 3.1 Pro in February).

> why_it_matters

The competitive order at the frontier shapes which models builders standardize on. A refocused DeepMind pushing hard on coding could reshuffle the leaderboard — and pricing — across the tools your team relies on.

> real_life_use_case

Avoid locking your stack to a single model provider — keep an abstraction layer so you can switch as the frontier leader shifts between Gemini, GPT and Claude over the next year.

> source

CNBC ↗

#07 · Policy

Musk proposes AI labs peer-review each other's models

# A rival-tests-rival idea enters the safety debate.
> the_news

Amid the industry-wide safety conversation sparked by Amodei's slow-down essay, Elon Musk called for AI peer review — proposing that competing labs test one another's models. The idea sits alongside OpenAI, Anthropic and Google's talks on a standards body and Hassabis's earlier FINRA-style proposal. Reporting notes coordination among competitors raises potential antitrust questions, though OpenAI's policy chief said the labs don't believe a waiver is needed. No formal cross-lab review program has been established yet.

> why_it_matters

Cross-lab evaluation would be a major shift in how model safety is validated — moving from self-assessment toward adversarial, third-party testing. It also reveals how fluid the governance debate still is, with proposals outpacing any actual rules.

> real_life_use_case

Teams in regulated sectors should document their own model-evaluation process now — if industry norms move toward peer review, being able to show a rigorous internal eval will become a procurement advantage.

#08 · Funding

Euno raises $23M to give AI agents enterprise “context”

# The unglamorous layer agents need: knowing what your data means.
> the_news

Euno raised a $23M Series A (total funding $29M) to expand its enterprise AI context platform, which gives agents the organizational knowledge to understand what corporate data means, which sources can be trusted, how data is used and what governance rules control access. The San Francisco/Tel Aviv company is betting that this “context layer” becomes its own infrastructure category as enterprises move agents from isolated assistants to workflows spanning many systems.

> why_it_matters

Agents fail in the enterprise less from weak reasoning than from missing context — they don't know which data is authoritative or permitted. Solving this is a precondition for agents that act correctly across Salesforce, Workday, email and more.

> real_life_use_case

Before scaling agents across departments, map your authoritative data sources and access rules — a clear context/governance layer is what separates a demo from a dependable production agent.

> source

citybiz ↗

#09 · Robotics

Skild AI unveils S1, a robot that learns 10-minute tasks from a single video

# A robotics foundation model that generalizes from one demonstration.
> the_news

Skild AI released S1, a robotics foundation model that executes tasks up to 10 minutes long from a single human video prompt — with no fine-tuning. The company reports 66% success on unseen tasks versus 9% for language-prompted vision-language-action models at the same 100k-hour training scale. It's part of a broader hardware-and-robotics wave that also saw Apple's first 2nm M6 chip and OpenAI's Jalapeño inference silicon debut in the same cycle.

> why_it_matters

“Show, don't fine-tune” is a big step toward general-purpose robots that ordinary operators can teach by demonstration. The jump from 9% to 66% on unseen tasks suggests robotics is approaching the generalization curve LLMs hit years ago.

> real_life_use_case

Operations and logistics teams should start cataloguing repetitive physical tasks that could be taught by a short video demo — that's the shape of work these models will target first.

#10 · Hardware

Apple debuts its first 2nm M6 chip, pushing on-device AI

# Local AI gets a meaningful hardware bump.
> the_news

Apple launched the M6 — its first 2-nanometer chip (12-core CPU, 12-core GPU, dual 16-core Neural Engine, up to 32GB unified memory at 170 GB/s) — debuting in a new Mac mini, alongside the quad-die M5 Ultra in the Mac Studio (up to 512GB memory at 1.2 TB/s). Apple claims the M6 delivers ~30% more peak GPU AI compute than M5, and reporting notes the new Mac mini runs up to ~4× faster AI performance, with the M5 Pro supporting up to 64GB for larger local models.

> why_it_matters

More unified memory and faster neural compute make it practical to run capable models fully on-device — no API bill, no data leaving the machine. That matters for privacy-sensitive and offline AI workloads.

> real_life_use_case

Developers building privacy-first tools can now target strong local inference on consumer Macs — worth prototyping an on-device model path as a premium, no-cloud option for sensitive users.

> source

PupuWeb ↗

# get this in your pocket every morning
Join on WhatsApp →
# AI Majlis Daily · Abu Dhabi · Dubai · Sharjah