MAI-Transcribe-2-Streaming provides real-time transcription across 60 languages with 100ms partial hypothesis latency, priced at $0.54 per audio hour through year-end.
MAI-Voice-2.1 supports 23 languages and 26 locales while preserving a single speaker's voice profile across languages for $22 per 1M characters.
MAI-Voice-2.1-Flash generates 45 seconds of audio with 150ms end-to-end latency and costs $15 per 1M characters.
Both voice models support voice cloning using a few seconds of reference audio with built-in consent guardrails.
A live demonstration application called Chatter is available in the MAI Playground.
The feature integrates TradingView Lightweight Charts to display financial data including candlesticks, volume, and moving averages directly inside threads.
ElevenLabs established an office in the Netherlands and plans to triple its local GTM and engineering headcount this year.
Goswijn Thijssen was appointed to lead the Netherlands operations.
The company fine-tuned its Dutch transcription models with local data, increasing Dutch postcode recognition accuracy from 64% to 82% for telecom client KPN.
Dutch telecom company KPN runs voice verification agents on the stack handling roughly 60,000 calls per week at a 62% success rate.
The ElevenLabs platform currently offers 100 Dutch voices.
OpenAI attributed a core cluster of the extraction attempts involving over 15,000 accounts to individuals associated with Moonshot AI, the developer of Kimi.
Operators manipulated model interactions and replayed encrypted reasoning across conversations to coerce models into decrypting hidden reasoning.
The campaign peaked on July 24 and 25 with 16,000 attempted extraction requests from 4,000 users before full disruption on July 28.
OpenAI deployed mitigations including banning accounts, closing reasoning replay pathways, and adding checks to hold streamed output exposing reasoning.
The company shared threat findings with industry partners through the Frontier Model Forum and government channels.
Skills let users save reusable instructions with reference files, automatically stack them for complex tasks, and will fully replace Gems starting in November.
Skills can include reference files such as plain text documents, PDFs, or images.
Multiple skills can be stacked and applied simultaneously to a single task.
Gemini can generate skills from chat histories and automatically invoke them when a prompt matches.
Gems will automatically migrate into skills starting in November for personal accounts and in 2027 for Workspace tiers.
Skills are rolling out globally today, with Google Workspace enterprise and education rollout following in coming weeks.
The watermarking tool embeds verification signals directly into AI-designed sequences to help DNA synthesis providers screen orders and track biological provenance.
Google DeepMind is open-sourcing the code, weights, and in vitro data alongside a methods paper.
Integrated with Evo 2 in collaboration with Stanford and the Arc Institute to generate functional watermarked bacteriophages verified in lab cultures.
Designed to flag synthetic submissions to public databases like GenBank, UniProt, and the Protein Data Bank.
The update adds bidirectional voice interaction, session forking into managed worktrees, and tools to monitor parallel agent tasks inside the terminal.
Supports two-way voice interaction directly in the terminal via the /voice command
Adds /fork to branch conversations into isolated Git worktrees with shared context
Introduces /agents to track and switch between concurrent background tasks
Renders Mermaid diagrams and LaTeX natively alongside collapsible diffs and tool outputs
Available by updating through npm with @openai/codex@latest
The GPT-6 Astra-powered agents run on dedicated cloud computers, connect with over 4,000 apps, and are rolling out to Pro, Business Premium, and Enterprise subscribers.
Each dot operates its own cloud computer and browser, with optional permission to access local laptops.
Users can message or voice-call dots through ChatGPT, Slack, and Microsoft Teams.
One dot is included at no extra cost for Pro and Business Premium accounts, with deep work allowances.
Proactive background research uses read-only tools to prevent unintended edits or actions.
OpenAI is piloting organizational specialist dots with dedicated credentials and Microsoft Agent 365 support.
The integration brings Cognition's Devin AI agent into MongoDB's Application Modernization Platform to automate rewriting legacy code, queries, and data access layers.
Devin rewrites business logic and data access layers while AMP migrates and validates data into MongoDB Atlas.
In early joint testing, migration tasks that took five to six hours took just over an hour.
The integration is available immediately for joint MongoDB and Cognition customers.
The model provides speech synthesis with a median inference latency of about 100 ms across more than 90 languages at 3.3 cents per minute through October 12th.
Agents adjust tone, emotion, pacing, and vocal delivery mid-call based on conversational context
Supports real-time language switching during live calls while maintaining voice consistency
Integrates alongside ElevenLabs transcription and turn-taking models to manage conversational interruptions
Includes pronunciation dictionary support for brand and technical terms across languages
The hub houses dedicated research teams collaborating with BMW on crash simulations, Siemens Energy on industrial AI, and TUM on aerodynamics.
Mistral aims to build one gigawatt of European compute capacity by 2030.
The hub integrates over 30 physicists and engineers from Mistral's May 2026 acquisition of Emmi AI to model computational fluid dynamics and multi-physics simulations.
A partnership with Technical University Munich utilizes wind tunnel facilities to fuse sensor data with physics simulations for real-time automotive aerodynamic predictions.
The models support over 90 languages and inline performance tags, with the Turbo variant delivering approximately 100 ms median inference latency.
Inline tags allow users to direct emotion, pacing, style, sound effects, and pronunciation via IPA support.
Instant Voice Clones generate voice models from 10 seconds of sample audio.
Language coverage spans 90+ languages, including Cantonese, Mongolian, and Odia.
Eleven v4 is available across ElevenAgents, ElevenCreative, and ElevenAPI, with temporary promotional API pricing of $22 per 1M characters for v4 and $11 for v4 Turbo.
Former MongoDB CEO CJ Desai joins Meta as Chief Enterprise Platform Officer to lead the new enterprise business offering tools like Muse API and Muse Code.
The enterprise platform will offer business tools including the Muse agent, Meta Business Agent, Muse API, and Muse Code.
Desai will report directly to Meta founder and CEO Mark Zuckerberg.
Desai previously served as CEO of MongoDB, President and COO at ServiceNow, and led product and engineering at Cloudflare.
The open-weight 7B World Action Model jointly predicts future video and robot actions, achieving up to a 42.2% success rate on RoboLab.
Single-step distillation achieves a 38.3% success rate on RoboLab while running 1.34x to 2.28x faster per second of motion than Pi0.5 in FP8 on workstation and datacenter GPUs.
Guidance-distilled checkpoints achieve 42.2% on RoboLab, outperforming Cosmos 3 Nano (36.8%) while delivering 1.52x to 3.95x speedups in FP8.
The model predicts a 2.13-second motion horizon compared to 1.0s for Vision-Language-Action baselines like Pi0.5.
Hybrid delegation with GPT 6 Astra achieves a 90% success rate on complex tasks while cutting costs to $8.77 and runtimes to 8 minutes per success.
Pretrained on video, image, and audio data, with midtraining across gaming inputs, egocentric hand poses, and teleoperation datasets.
The new tier runs on the custom Rust engine Photon, returning 95% of results within 230 ms and cutting agent task costs by 68%.
Median latency is 160 ms by Perplexity's benchmarks.
Fast Search scores 0.24 points lower on relevance and 3 percentage points lower on answer availability on internal long-tail tests compared to the default preset.
Photon replaces Perplexity's previous open-source retrieval engine across the entire service, lowering internal p99 response times from 800 ms to 65 ms.
Photon uses 20% fewer serving machines, stores 2.5x as much data per document, and isolates index builds from live query servers.
The feature combines real-time video generation and speech in Gemini Enterprise, providing synchronized lip-syncing and expressions across 97 languages.
Supports asynchronous tool calling, fetching data in the background while maintaining continuous dialogue
Adapts expressions and lip-syncing across 97 languages without visual drift
Generates custom responsive avatars from a single reference image via enterprise allowlisting
Watermarks all audio and video output using SynthID
The expansion covers Indonesia, Malaysia, the Philippines, Singapore, Thailand, Vietnam, and Taiwan, bringing total ChatGPT Ads availability to more than 60 countries.
Ads are shown only to Free and Go plan users, while Plus, Pro, and Enterprise tiers remain ad-free.
Eligible businesses can purchase inventory through OpenAI's self-service Ads Manager or agency partners including dentsu, Havas Media, Omnicom Media, Publicis Groupe, and WPP.
OpenAI reports ChatGPT Ads surpassed a $1 billion annualized revenue run rate in under 200 days post-launch.
The smart glasses add voice-driven digital avatars for video calls, landmark-based navigation, Dolby Atmos capture, and online ordering across five new countries.
Hologram generates a voice-driven photorealistic face avatar for WhatsApp video calls using speech inference after a mobile app capture.
Navigation adds cycling, public transit departures, and visual landmark prompts rather than street names.
Video recording now captures spatial surround sound natively in Dolby Atmos.
Online ordering and retail availability expand to Canada, the UK, France, Italy, and Germany.
The feature executes tasks on Anthropic-hosted infrastructure without requiring a local machine to remain open, offering one-time credits of $100 for Pro and $250 for Max users.
Sessions can be started at claude.ai/code, through desktop and mobile apps, or using claude --cloud in the CLI.
A connected GitHub account is required to start a session.
Promotional credits are separate from regular usage limits and must be claimed by Oct 7.
The two text-to-speech models offer prompt-driven voice design across 100+ languages, 30-second voice replication with verbal consent checks, and line-by-line performance direction.
Gemini 3.8 Flash TTS focuses on creative character direction, while Flash-Lite TTS is optimized for cost-efficient, high-volume dubbing and voice agents.
Features include two-speaker script staging, long-form generation with minimal drift, and scripted non-verbal cues like laughs and gasps.
Voice replication requires a matching verbal consent recording and embeds SynthID watermarks alongside C2PA credentials in generated audio.
Available immediately in Google AI Studio and the Gemini API, alongside integrations in Gemini Notebook and Google Vids.
The method trains models on real-world sessions by using corrective hints during training to align hint-free next-token predictions with hint-guided outputs.
In live A/B testing, a checkpoint trained with the technique reduced tool-call failures by 21.2% relative to an earlier checkpoint, dropping failure rates from 2.24% to 1.77% without hints at inference.
The annotation pipeline traces feedback to specific decisions and validates corrective hints against information available before the mistake to minimize hindsight bias.
GLM 5.2 scores recorded turns twice using the same weights—with and without hints—aligning unassisted predictions with guided predictions.
Providing hints enabled the unchanged model to avoid original failures in 93.7% of cases, up from 75.1% unassisted.
Training samples filter out personally identifiable information and exclude sessions from opted-out users.
Anthropic released Claude Opus 5.5 alongside a showcase of early community explorations spanning interactive interfaces, simulations, and generative tools.
Demos include a pen-and-paper themed Claude Code interface where new sessions open as napkins
Showcased applications include a toy brick-building tool that converts descriptions and photos into buildable instructions
Other examples include an algorithmic seed-based drawing tool and an interface slot machine for generating random app layouts
The model cuts clinical word error rates by 35% compared to base Scribe v2 and is available via API starting at $0.22 per hour.
Achieved the lowest overall WER on the MedDictate clinical dictation benchmark across English, French, and German.
Scored lower WER and CER than published leaderboard results on the 3,619-sample Eka Medical ASR benchmark.
Maintains a 5.3% WER on 6,000 non-medical audio samples from Common Voice, matching base Scribe v2.
Enterprise customers with a BAA can enable Zero Retention Mode to delete audio and transcripts immediately upon request completion for HIPAA compliance.
The new model supports French, German, Hindi, Italian, Polish, Portuguese, and Spanish across all subscription tiers on the ElevenLabs beta platform.
Compatible with VoiceLab features including Instant Voice Cloning and Voice Design while maintaining speaker voice characteristics and accents across languages.
Generates speech across multiple languages from a single prompt while preserving speaker identity.
Known limitations include numbers, acronyms, and foreign words occasionally defaulting to English pronunciation when prompted in other languages.
Hosted at the Institute for Advanced Study, the independent, unpaid group will advise OpenAI on research standards and the dissemination of math capabilities.
The advisory group operates independently, can offer unprompted feedback, publish commentary publicly, and adjust its own membership.
Members are unpaid by OpenAI and will not advise on the pacing of OpenAI's internal model progress.
Initial members include Timothy Gowers, Martin Hairer, Edward Witten, Melanie Matchett Wood, and Camillo De Lellis.
OpenAI reported that a new internal model solved the Navier–Stokes Millennium Prize problem and over 100 open math problems.
The platform integrates multimodal generation directly into a rebuilt multi-track timeline alongside Studio Agent, an AI co-editor that executes edits from text prompts.
Studio Agent analyzes footage frame by frame to draft first cuts, sync audio cues, and apply edits from natural language descriptions
Generates video, images, music, sound effects, and voiceovers across more than 10,000 voices in over 32 languages
The timeline includes frame-level zooming, clip snapping, synced clip movement, in-timeline caption editing, and a rebuilt playback engine
Adds team collaboration features with timestamped and clip-level commenting directly on project assets
The standalone desktop application includes multi-agent parallel execution, an integrated browser and terminal, and support for third-party model providers.
Supports multi-agent parallel workflows and background execution for long-horizon programming tasks
Integrates a side-panel browser, terminal, code diff viewer, and file preview directly within the application
Allows users to highlight webpage elements, annotate screenshots, and drag-and-drop folders into the prompt input
Provides three agent permission tiers: always ask, ask when needed, and fully automatic execution
Features a plugin marketplace and allows developers to bring their own model providers