Kamil Józwik

Developer AI news

July 2026

actAVA introduces CURA 1T, a 1-trillion parameter clinical AI model

Healthcare AI company actAVA launched CURA 1T, a 1-trillion parameter clinical model. It claims to outperform frontier rivals on healthcare benchmarks at significantly lower costs, offering a specialized and efficient solution for enterprise healthcare applications.

PrismML releases Bonsai 27B, an iPhone-compatible small model

PrismML launched Bonsai 27B, a 27-billion parameter model designed to be small enough to run efficiently on iPhones. This release signifies progress in making powerful AI models accessible for on-device mobile deployment, enabling advanced AI capabilities on consumer hardware.

OpenAI publishes research on GPT-Red, an automated safety red-teaming model

OpenAI released new research on GPT-Red, an internal automated safety red-teaming model used to train GPT-5.6 against prompt injections. This highlights technical efforts in AI safety and robustness for frontier models, aiming to improve their resilience to adversarial inputs.

Weco's AIDE² agent demonstrates recursive self-improvement in research

AI research team Weco demonstrated early recursive self-improvement with AIDE², an automated agent that redesigned its own research process over eight days. The system tested 100 rewrites, keeping seven upgrades, and ultimately outperformed a hand-built version on benchmarks, showcasing potential for advanced AI capabilities.

Thinking Machines Lab releases Inkling, its first open-weights multimodal model

Mira Murati's Thinking Machines Lab introduced Inkling, its first open-weights multimodal model, designed for customization and fine-tuning rather than raw benchmark scores. Strong in agentic web design, instruction following, and math reasoning, Inkling can be downloaded via Hugging Face and customized through TML's Tinker service.

Roblox unveils 'Build' AI tab for generating games from text prompts

Roblox introduced 'Build,' a new AI tab for its mobile app that allows users to generate playable games directly from text prompts. This feature aims to democratize game creation and expand the platform's developer tools, making game development more accessible.

Google rebrands NotebookLM to Gemini Notebook, adds cloud computer tool

Google rebranded NotebookLM as Gemini Notebook, integrating a new cloud computer tool to enable code execution for enhanced data analysis. This update aims to improve the platform's capabilities for developers and data scientists, offering more robust analytical features.

Guide: Use OpenAI's GPT-Live for efficient trip planning with voice AI

A guide demonstrates using OpenAI's GPT-Live (Voice AI) on desktop or mobile to plan detailed trips. Users can direct the AI, interrupt for changes, and leverage web research for verification, then export the plan as a memo, with a pro tip to set an iPhone Action Button for quick access.

Moonshot AI releases Kimi K3, an open-weights model competitive with frontier LLMs

Moonshot AI launched Kimi K3, an open-weights model featuring a 1M context window that outperforms Claude Fable 5 and GPT-5.6 Sol on benchmarks for web research, spreadsheet work, frontend design, and long coding. Priced at $3/$15 per million tokens, its weights are slated for release by July 27.

PicLumen offers free AI image/video generation with commercial license

PicLumen is a free AI image and video generator that provides unlimited daily generations on a relaxed queue, multiple models (photoreal, anime, FLUX), and a built-in editor. Crucially, its free tier includes a commercial license, making it a valuable resource for developers needing commercial-safe assets.

Report: OpenAI's first device is a screen-free, camera-equipped AI speaker

Bloomberg reports OpenAI's Jony Ive-designed hardware will be a screen-free, battery-powered speaker with cameras and sensors, capable of controlling smart-home devices and personalizing interactions. This device, expected in 2027, aims to establish a new AI computing interface for the home.

Anthropic's Claude Opus 5 spotted on Google Vertex, expected soon

Claude-Opus-5 was observed on Google Vertex, signaling an imminent release, likely within the next week. This new model is anticipated to be significantly cheaper and more token-efficient than Fable 5, positioning Anthropic competitively in the ongoing AI model price war.

OpenAI increases ChatGPT custom instructions limit to 5,000 characters

OpenAI has raised the custom instructions limit for ChatGPT from 1,500 to 5,000 characters for Pro, Business, Enterprise, and Education users. This update allows for more detailed and persistent guidance for AI interactions, enhancing customization capabilities for developers.

Manus AI now generates editable PowerPoint files directly

Manus, an AI tool, has been updated to generate PowerPoint files directly, preserving chart editability and allowing real-time value changes and element toggling. This enhancement streamlines presentation creation, offering a more flexible workflow for users.

SpaceXAI open-sources Grok Build coding tool after privacy concerns

SpaceXAI open-sourced its Grok Build CLI, a coding agent, and is deleting all retained coding data after a researcher discovered it was uploading entire private repositories to a cloud bucket. This move aims to enhance reliability and directly address privacy questions raised by the beta version's data defaults.

Hacker leaks Suno's training data scraping methods from YouTube, Deezer, and Genius

A hacker breached AI music generator Suno, leaking source code that details scraping routines which pulled 2 million clips from YouTube Music, 1 million hours from podcasts, and data from Deezer and Genius. The code specifically targeted acapella versions and utilized Bright Data proxies to bypass YouTube's defenses, validating long-standing copyright concerns.

Google delays Gemini 3.5 Pro launch due to performance issues

Google postponed the release of Gemini 3.5 Pro after internal testing revealed newer builds performed worse than older checkpoints, jeopardizing its July 17 target. This delay underscores challenges in maintaining model performance during development, with internal conflicts and employee departures cited as contributing factors.

NVIDIA Build platform for testing frontier and open AI models

NVIDIA Build is a platform enabling developers to test frontier and open AI models directly in a browser, eliminating the need for installation or GPU rental. It offers a playground for experimenting with various models, including Physical AI and Agentic AI, and provides API snippets for easy integration into projects.

OpenAI launches Codex Micro, a $230 keypad for AI agent control

OpenAI released its first branded hardware, the Codex Micro, a $230 keypad co-developed with Work Louder for controlling AI coding agents. It features 'Agent Keys' for task color-coding, a joystick for actions like code reviews, and a dial for adjusting agent reasoning levels, providing a physical interface for agentic workflows.

Nvidia unveils Cosmos 3 Edge, an AI model for robots and vision agents

Nvidia launched Cosmos 3 Edge, a world model designed for robots and vision agents, capable of real-time physical space perception and movement from diverse inputs. Running on-device with Nvidia's Jetson Thor boards, it enables robots to reason without cloud round trips, expanding Nvidia's AI ecosystem into the physical world.

ByteDance launches Seedream 5.0 Pro image model for advanced design work

ByteDance has rolled out Seedream 5.0 Pro, a new image model designed to "understand design" with powerful editing features, improved text rendering, and layer separation for editable designs. It aims to provide frontier-level control and creative partnership for user workflows.

ViewGen: Open-source Unreal Engine plugin integrating ComfyUI for 3D viewport data

Michael N. developed ViewGen, an open-source Unreal Engine plugin that integrates ComfyUI to leverage 3D viewport and camera data. This allows users to generate precise depth and color reference images and movies from UE, piping them into Comfy-styled graphs directly within the engine for advanced 3D development.

Guide: Enhance employee 1-on-1s with Claude Cowork for structured feedback

This guide demonstrates how to leverage Claude Cowork to improve employee 1-on-1s by setting up a project with past transcripts, creating a customized call template, and an evaluation rubric. It enables managers to track progress and automate template updates for consistent feedback.

AI workflow for commercial mortgage identification and prospect ranking using Claude

A commercial loan portfolio manager developed an AI workflow using Claude to process public mortgage data, extract commercial filings, enrich contact information, and rank prospects by priority. The system generates an Excel file with an executive summary and targeted suggestions for outreach.

Anthropic introduces 'Reflections' dashboard for Claude usage analysis and optimization

Anthropic has launched a 'Reflections' dashboard to help users analyze their interaction patterns with Claude. This feature provides insights into usage habits and suggests skill creation based on personal data, aiming to optimize user experience and prompt engineering.

Meta to begin manufacturing in-house 'Iris' AI chip in September

Meta is reportedly set to begin manufacturing its proprietary 'Iris' AI chip in September, with plans to double its computing capacity to 14 GW by 2027. This move signifies Meta's increasing investment in custom AI hardware to support its growing AI infrastructure.

RobbyAnt releases LingBot-World 2, a real-time world model for embodied AI

Embodied AI lab RobbyAnt has launched LingBot-World 2, a new world model capable of generating every frame in real-time without requiring a 3D engine. This represents a significant advancement for embodied AI systems, enabling more dynamic and responsive virtual environments.

OpenAI retracts endorsement of SWE-Bench Pro coding benchmark due to task issues

OpenAI has published research retracting its endorsement of the SWE-Bench Pro coding benchmark, after discovering that nearly a third of the benchmark's tasks contained issues. This highlights challenges in accurately evaluating AI coding capabilities and the need for robust benchmarks.

Reve 2.1 image model released, achieving high Arena ranking with less compute

Reve has launched version 2.1 of its native-4K image model, which has re-secured the No. 2 spot on Arena's overall leaderboard. This model achieves its high performance while being trained on less than a tenth of the compute used by its rivals, highlighting significant efficiency gains.

Guide: Optimize Fable token usage with an orchestrator setup for cost efficiency

This guide details how to reduce Fable token consumption by strategically using it as a planner and reviewer, while delegating token-heavy tasks like browsing, coding, and research to lower-cost models such as Codex or other Claude variants. It provides a step-by-step orchestrator workflow.

Microsoft replaces third-party AI models with in-house MAI in Office applications

Microsoft is transitioning tens of thousands of prompts in Excel and Outlook to its proprietary MAI models, aiming to reduce and ultimately eliminate reliance on external AI providers like OpenAI and Anthropic. This strategic shift enhances internal control and optimizes costs.

Anthropic's Claude Cowork expands to mobile and web with background task execution

Claude Cowork, Anthropic's agent, is now available on mobile and web, allowing tasks to run in the background even when devices are offline. This enables seamless project continuity across devices, though desktop remains exclusive for local file and browser access.

Cognition releases SWE-1.7, a fast, near-frontier coding model for Devin

Cognition has launched SWE-1.7, a new coding model built on a Kimi K2.7 base, which operates at 1,000 tokens per second. It achieves performance comparable to frontier models and is integrated into Devin, served on Cerebras, enhancing automated code generation.

Willow offers free, unlimited AI dictation with Frontier Mini model

Willow has launched its Frontier Mini model, providing free and unlimited cloud-based AI dictation with zero data retention. It claims superior speed and accuracy compared to competitors like Wispr Flow, OpenAI, and Deepgram for transcribing speech into text across applications.

OpenAI introduces GPT-Live for natural, full-duplex voice conversations

GPT-Live, OpenAI's next-gen voice model, replaces ChatGPT Voice with a full-duplex architecture that enables simultaneous listening and talking, eliminating awkward pauses. It can delegate complex reasoning to stronger models and supports real-time translation, offering a more human-like conversational experience.

Google DeepMind's AlphaEvolve algorithm-inventing agent now generally available

AlphaEvolve, the DeepMind system capable of writing and improving its own code, has reached general availability. Klarna has already utilized it to double its machine-learning training throughput, demonstrating its utility for optimizing ML training processes.

Google Cloud Run sandboxes enter public preview for secure AI code execution

Google Cloud Run sandboxes are now in public preview, offering isolated, ephemeral environments for AI agents to securely run code. These sandboxes provide zero network access by default, spin up in milliseconds, and are available at no additional cost beyond standard Cloud Run fees.

Mistral Releases Leanstral 1.5 Open Model for Verified Math Proofs

Mistral introduced Leanstral 1.5, an open model specifically designed for the generation and verification of mathematical proofs. This model aims to advance automated reasoning and formal verification in mathematics.

ByteDance Introduces Seed Audio 1.0 for Unified Audio Generation

ByteDance launched Seed Audio 1.0, a versatile model capable of generating speech, music, and sound effects in a single pass. This model offers a unified solution for various audio content creation needs.

Kyutai and General Intuition Release MIRA World Model for Rocket League Simulation

Kyutai and General Intuition, in collaboration with Epic Games, released MIRA, an open-source world model capable of running a live 2v2 Rocket League game entirely within a neural network. Trained on bot footage, it simulates game physics and renders details at 20 FPS on a single Nvidia GPU.

Replit Enables Rapid Mobile App Prototyping with AI Agent

Replit provides tools and a guide for developers to rapidly prototype mobile applications using its AI Agent. This allows quick transformation of app ideas into testable prototypes with Expo Go, facilitating iterative refinement of core user flows.

Tencent Hunyuan Open-Sources Efficient Hy3 Model with Apache 2.0 License

Tencent's Hunyuan released Hy3 as an open-source model under an Apache 2.0 license, achieving high efficiency by using a small subset of parameters per request. It competes with larger models on web research and tool use, offering a compelling option for developers.

Anthropic Research Uncovers 'J-space' Internal Workspace in Claude

Anthropic's research revealed 'J-space,' an internal workspace within Claude that functions like an internal notepad, holding active concepts and directing the model's thinking. This undesigned structure emerged during training and is crucial for multi-step problem completion.

Tufa Labs Wins ARC-AGI-3 Contest with Open-Source Qwen-Based Coding Agent

Tufa Labs secured first place in the ARC-AGI-3 milestone contest by developing an open-sourced coding agent. This agent, which wraps a small Qwen model, demonstrated effective problem-solving capabilities on game-based benchmarks.

DoorDash Introduces DashBench for Evaluating AI Code Reviewers

DoorDash developed DashBench, an internal benchmark to test AI code reviewers against historical code changes. A pairing of Kimi K2.6 and Claude Fable 5 achieved the best results, catching two-thirds of problems and 80% of critical bugs at $3.81 per review, enabling open models in their pipeline.

Google for Startups Releases Generative Media Technical Guide

Google for Startups launched a technical guide for building production-grade, multimodal creative AI applications using DeepMind models like Veo and Lyria. The blueprint emphasizes deterministic control, programmatic guardrails, cryptographic provenance, and economic scalability.

Researchers Develop 23MB Add-on for Tiny Models to Match Qwen3-32B Performance

Researchers have created a 23MB add-on that enables small AI models to achieve performance comparable to Qwen3-32B on common text tasks. This innovation allows for efficient offline execution on devices like a MacBook, significantly reducing model size requirements.

Cassidy Launches No-Code AI Agent Builder for Enterprise Workflows

Cassidy provides a no-code platform for building model-agnostic AI agents that integrate with enterprise tools like CRMs and chat apps. These agents automate repetitive tasks such as drafting proposals and triaging tickets, deployable in Slack, Teams, or browsers.

Meta's 'Watermelon' Model Reportedly Matches GPT-5.5 Performance

Meta's 'Watermelon' model, currently in training, is reported to achieve performance on par with OpenAI's GPT-5.5, albeit using 10x the compute of Muse Spark. An update to Muse Spark with significant coding and agentic improvements is also planned for Meta AI and its API.

Open-Source pxpipe Tool Reduces Claude API Costs by 70%

The open-source tool pxpipe enables developers to cut Claude Code API costs by up to 70% by converting text to PNG images, leveraging cheaper image token pricing. While it introduces some lossiness and latency, it offers significant savings for high-volume usage.