Comfy-Org/MiniMax-H3
MiniMax-H3 is trending on Hugging Face with over 12 million downloads. It's a multimodal model for image and text tasks, available to download and run locally.
Sherbert. Serving up the daily AI scoop.
MiniMaxAI open-sourced MiniMax-H3, an image-to-video model that passed 2 million Hugging Face downloads in days. Its Comfy-Org variant hit 12 million downloads while Qwen3.8, Kimi-K3, and DeepSeek-V4-Flash also trend hard on the leaderboard. Unsloth now runs them all locally and tools like MoneyPrinterTurbo turn any keyword into a finished HD short video without cloud bills.
What I clipped from today's AI news.
MiniMax-H3 is trending on Hugging Face with over 12 million downloads. It's a multimodal model for image and text tasks, available to download and run locally.
Kimi-K3 by Moonshot AI handles image-to-text and text tasks. Open weights, free to download, and hitting strong engagement on Hugging Face.
MoneyPrinterTurbo generates HD short videos from a topic or keyword using AI and automation. Open source, free to run locally.
Qwen3.8-27B is a multimodal model for image and text understanding. Open weights, free to download and deploy.
RAGFlow is an open-source retrieval-augmented generation engine that adds agent capabilities to improve context for LLMs. Free to self-host.
Unsloth is a local UI for running and fine-tuning LLMs and diffusion models including Qwen, Kimi, MiniMax, Gemma, DeepSeek, and FLUX. Free and open source.
MiniMax-H3 by MiniMaxAI generates video from images and text. Open weights, over 2 million downloads, free to use.
DeepSeek-V4-Flash is a lightweight text generation model trending on Hugging Face. Open weights, free to download.
Anthropic released a collection of Jupyter notebooks showing practical ways to use Claude. Free reference material for building with the API.
A fine-tuned variant of Qwen3.6-27B for image-to-text tasks, available as GGUF for local inference. Free to download.
Dive into LLMs is a Chinese-language programming tutorial series for learning large language models. Free educational resource.
PPT Master converts documents or topics into native PowerPoint decks with shapes, transitions, charts, and audio narration. Open source, free to run locally.
exo lets you run frontier models locally by distributing inference across multiple machines. Free to download and use.
ToolJet is an open-source platform for building internal tools, dashboards, and AI agents without writing much code. Free to download and self-host.
A community repository of prompts, agents, and configurations for GitHub Copilot. Free to use.
Scientific-agent-skills is a library of 161 pre-built skills for turning any AI agent into a research tool across biology, chemistry, medicine, and drug discovery. Free to download.
RAG_Techniques is a collection of advanced retrieval-augmented generation methods with detailed Jupyter notebook tutorials. Free to explore.
OpenViking is a context database for AI agents that unifies memory, knowledge retrieval, and skills in one system. Free to download.
SHAP is a library that explains why machine learning models make specific predictions using game theory. Free to use.
The fastai book, published as Jupyter Notebooks
SAM 2 is Meta's segmentation model for identifying objects in images and videos with high accuracy. Free to download and use.
Unsloth released a quantized version of Qwen 3.8 27B optimized for faster inference on consumer hardware. Free to download.
Muse-Glimmer-30B is an open model that generates text from images and text prompts. Free to download.
A set of 29 editorial diagram templates for Claude Code as clean HTML and SVG, no bloat. Free to use.
Google Cloud released sample code and notebooks for building with Gemini on their platform, including templates for the Enterprise Agent framework. Free to download and use.
Chip Huyen published an open AI engineering resource book with supporting materials and code examples. Work in progress, free to read.
Microsoft open-sourced a curriculum for Model Context Protocol (MCP), the standard for connecting AI agents to external tools and data. Covers .NET, Java, TypeScript, JavaScript, Rust, and Python with real-world examples.
Holehe is an open tool that checks whether an email address is registered on sites like Twitter and Instagram, and retrieves account info from password recovery endpoints. Free to download and run.
Lightricks released LTX-2.5, an open image-to-video model available on Hugging Face. Free to download and run locally.
NVIDIA open-sourced Cosmos, a platform of world models and datasets for building physical AI in robotics, autonomous vehicles, and smart infrastructure. Free to use.
Unsloth released Muse-Glimmer-30B, a quantized open model that takes images and text as input and generates text. Free to download.
Ego-lite is a lightweight browser built for AI agents to run automation tasks while keeping your logged-in sessions intact. Free and open-source.
Alibaba released Qwen3.8-2.4T-A95B, a large open text generation model. Free to download.
Lightricks published the official Python package for LTX-2, their audio-video generative model, with inference and LoRA fine-tuning support. Free to download.
Semantica released a graph-native infrastructure framework for building AI systems with better context handling and accountability. Open-source.
Liquid AI released LFM2.5-2.6B, a small open text generation model. Free to download.
MiniMax-H3-Turbo-Lora is a text-to-video model fine-tuned for faster inference. Free to download and run locally.
holaOS is an open-source agent workspace that chains Claude, Codex, and other agents across 100+ integrations, MCP servers, browser, and files with shared memory. Free to download and self-host.
MiniMax-Music3 generates audio from text prompts. Free to download and run locally.
Minimax-H3-Turbo converts images to video. Free to download and run locally. Over 211k downloads suggests solid adoption.
ai_quant_trade is an end-to-end platform for stock trading that combines LLMs, machine learning, reinforcement learning, and high-frequency trading strategies. Free to download and self-host.
Modly is a desktop app that generates 3D models from images or text prompts using local AI on your GPU. Free to download.
Complete Agentic AI Engineering Course repository with code examples and tutorials. Free to access.
Muse-Glimmer-30B is a multimodal model that takes images and text to generate text responses. Free to download and run locally.
DeepSeek's awesome-deepseek-agent is a curated collection of agent frameworks and examples built on DeepSeek models. Free to access.
Qwen3.8-27B-FP8 is a quantized multimodal model that processes images and text. Free to download and run locally.
DeepSeek-V4-Pro-0813 is a text generation model released on Hugging Face. Free to download and run locally.
Oracle AI Developer Hub provides technical resources for building AI applications and agents on Oracle's database and cloud infrastructure. Free to access.
NVIDIA released Nemotron-3.5-Lightning, a 30B text model optimized for speed and inference cost. Open weights, free to download.
Kijai ported MiniMax-H3 to ComfyUI, the node-based image and video generation interface. Free to download.
A fine-tuned MiniMax-H3 text-to-video model is circulating on Hugging Face with a focus on realism. Free to download.
Macro is an open-source workspace that bundles email, chat, docs, tasks, and agents with shared AI memory across tools. Free to download and self-host.
Ling-3.0-tiny is a small language model from inclusionAI, trending on Hugging Face. Free to download.
A LoRA adapter for MiniMax-H3 tuned toward photorealistic people in video generation is available on Hugging Face. Free to download.
Qwen released Qwen3.8-2.4T, a large text model quantized to 8-bit for smaller footprint. Open weights, free to download.
Anima-2.9B is a small text-to-image model gaining traction on Hugging Face. Free to download.
Qwen 3.8 27B landed and drew heavy discussion on Hacker News. Closed weights, API pricing not yet announced.
Firefox is now the only major browser still allowing uBlock Origin to run without restrictions. Chrome, Edge, and Safari have moved to Manifest V3, which limits ad blocker power.
NVIDIA NeMo's Automodel is a PyTorch library for distributed training of large language and vision models with built-in Hugging Face integration. Free to download and use.
Australia's home battery boom is pushing wholesale power prices down, showing how distributed energy storage reshapes grid economics. Not AI, but worth watching for how it affects inference infrastructure costs.
Anthropic published a guide to getting more out of Claude's code execution feature, which lets the model run and debug scripts in real time. Worth a look if you're using Claude for development work.
LuaCAD is a parametric CAD tool you script in Lua instead of clicking through menus. Open source and free to download.
Xiaomi's phone camera confused the moon for the sun during an eclipse, applying aggressive brightening to what it thought was underexposed. A reminder that AI image processing still trips on edge cases.
Mole is a terminal-based research agent that digs into topics and compiles findings. Free and open source.
Private prison operators reported $1.4 billion in revenue as immigration detention surged. Not AI news, but landed on HN and worth noting if you follow how AI shapes policy and enforcement.
Unsloth released quantized GGUF versions of Qwen 3.8 27B, making the model smaller and faster to run locally. Free to download.
Alibaba's Qwen 3.8 27B is out. It is a mid-size open weights model competitive with similar-scale peers. Free to download.
AI Model Atlas visualizes the landscape of ML models as an interactive 3D graph, showing relationships and lineage. Free to explore.
Anthropic published a research paper on AI safety risks and mitigations, dated August 2026. Signals their thinking on long-term alignment challenges.
Anthropic explained how Claude's text watermarking works: a hidden signal embedded in token choices that proves the text came from Claude. Technical deep dive on provenance.
The US military lost about one-fourth of its drone fleet in operations against Iran, straining supply chains. Not AI news directly, but relevant to autonomous systems and defense procurement.
YouTuber Davie504 received a copyright strike from an AI-generated music detector, raising questions about false positives and automated enforcement. Signals growing friction between AI tools and content platforms.
Researchers tightened the bounds on the Grothendieck constant, a fundamental problem in math that connects to neural network theory and optimization. Progress on hard constants like this often unlocks new approaches to AI training.
A new verifier checks whether LLM-generated GPU kernels (code that runs on graphics processors) are correct before deployment. This matters because AI models often generate code, and broken GPU kernels can crash inference.
Graft is a Claude integration that hooks into your code editor and cuts the number of tokens sent to the API by 42 percent through smarter context selection. Free to try.
Open WireGuard Endpoints is a tool for setting up encrypted network tunnels without a central server. Not AI-specific, but relevant to anyone running local models or private inference.
This is a news story about geopolitics, not AI. It does not belong in Sherbert.
ThoughtDAG is a tool that lets you edit the context graph inside an LLM conversation, so you can prune, reorder, or add reasoning steps before the model responds. Free to try.
Alibaba released Qwen 3.8 27B, an open-weight model on Hugging Face. At 27 billion parameters, it competes with mid-range closed models and runs on consumer hardware.
MattVidPro benchmarked Qwen 3.8 27B against other local models and found it competitive with much larger closed models like Claude Opus. Open-weight, free to download.
20 points, 3 comments
🚨 Everything you’re about to see, I benchmarked using the tool I made. If you want to run these tests yourself or test models on whatever you actually use them for! → https://www.woaibench.ai/ GLM 5.3 Is INSANE! The BEST Open Source Model EVER? (Fully Tested) Z.ai is back with GLM-5.3, and this could be one of the most impressive open-weight AI model releases we've seen yet. What's crazy is that GLM-5.3 uses the same base model as GLM-5.2 — the massive improvements come almost entirely from sca
OpenAI is launching "Ultrafast," a new inference mode that delivers GPT-5.6 Sol at up to 750 output tokens per second, powered by Cerebras hardware from their $10 billion partnership. Together with "Standard" and "Fast," Ultrafast creates a three-tier pricing structure that turns inference speed into its own product. The article GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras appeared first on The Decoder .
Alibaba's AI team Qwen has released new open model weights under the Apache 2.0 license with Qwen 3.8. The dense 27-billion-parameter model is designed to outperform the larger Qwen 3.7 Plus in coding and office tasks and natively processes up to 262,000 tokens of context. With this release, Qwen is targeting developers building local and agent-based applications. The article Alibaba's Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license appeared first on The Decoder .
AI agents using Claude Opus 4.8 and GPT-5.6 Sol were given six days, $3,000 in API credits, and GPU access to independently write AI research papers. The original authors of unpublished NeurIPS papers rated the results as "Reject." According to the study, conducted with Princeton and the UK AI Security Institute, frontier models can handle the full research engineering process but fall short on research judgment, creative problem-solving, and the ability to abandon failed approaches. The article Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach appeared
Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on r
Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage. The article New benchmark confirms AI models still perform poorly at visual perception appeared first on The Decoder .
Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the company’s more powerful model that stays locked behind its own APIs. The release landed alongside a letter from Mark Zuckerberg arguing AI should be “for everyone” rather than controlled by a handful of labs, but as Equity’s […]
Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the company’s more powerful model that stays locked behind its own APIs. The release landed alongside a letter from Mark Zuckerberg arguing AI should be “for everyone” rather than controlled by a handful of labs, but as Equity’s […]
Tim O’Reilly built a publishing empire that AI is helping to destroy. Yet he loves AI—as long as it’s open source.
Most people are barely scratching the surface of what the new Codex app can do. It can now browse social media without an API, turn something you build into a live, shareable website, and let you control work running on your computer directly from your phone. To try it for yourself, you can use Codex with both the free and paid plans, although free accounts have lower usage limits and some features have separate requirements. Download the ChatGPT desktop app with the link in the pinned comme
Score Distillation Sampling (SDS) enables text-to-3D generation by optimizing rendered images with a pretrained diffusion prior, but latent SDS often produces structured color artifacts and high-frequency texture noise. We identify a failure mode of latent SDS caused by VAE-induced pixel drift: the optimized image can move along pixel-space directions that are weakly constrained by the VAE encoder, so its latent representation remains clean and semantically meaningful while the image itself accumulates visible artifacts. We support this diagnosis with controlled 2D SDS experiments, VAE-only op
Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools, such as depth estimation and 3D reconstruction tools, to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning thro
This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-
Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruc
Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active interv
Anthropic will soon offer a watermark detection API that lets third parties check whether text was written by Claude. The technology builds on Google's SynthID method and tweaks the randomness during word selection without affecting text quality, Anthropic says. The approach has limits with fact-heavy text, code, and heavy rewriting. The article Anthropic announces watermark detection API that will let third parties detect Claude's AI texts appeared first on The Decoder .
The idea that GPUs are poorly suited for agentic workflows may be a misconception, according to French startup Kog.
Apple has reportedly trained a custom AI model for the China market alongside domestic tech giant Alibaba, a rare cross-border partnership that cuts across growing tensions between Beijing and Washington. The China-focused large language model was developed in partnership with Alibaba and trained with the company's support, Reuters reports, citing three unnamed people familiar with […]
❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here: https://www.anthropic.com/research/riemann-zeta Source: https://www.scientificamerican.com/article/no-ai-didnt-just-solve-the-thorniest-problem-in-math/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible: Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan
Z.ai released GLM-5.3 on August 14, 2026. The model reuses the 743B GLM-5.2 base unchanged. Every reported gain comes from scaled post-training: more long-horizon task environments, more environment types, longer training. Terminal-Bench 3.0 moves from 4.6 to 28.3, and DeepSWE v1.1 from 46.2 to 66.9. Cybersecurity moved further than Z.ai says it planned, with CyberGym at 84.5% and ExploitBench more than doubling to 54.4%. Weights arrive in about two weeks. The post Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks appeared first on MarkTechPo
Cactus Compute released Needle 2, an open 45M-parameter model for tool calling, device use, and structured extraction. The full model is a single 14MB binary that runs a session in about 28MB of RAM. It leads both Seal-Tools splits while targeting hardware with no GPU and no NPU. The post Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM appeared first on MarkTechPost .
Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding. AVA-Encoder transforms a video into a knowledge graph (KG) representation and then reconstructs it back into video. Its
Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real-world tasks. The rapid capability enhancement of open-source models deployable on consumer-grade GPUs presents a compe
Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from how much compute to spend, to where to allocate it. We formalize test-time reasoning as a constrained compute allocation problem over partial trajectories. Under a fixed hardware budget, existing paradigms fail to actively allocate the compute to the most promising partial progress: traditional parallel sampling treats traces independently and induces severe memory bottlenecks, while subtractive pruning starves ha
An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fi
Google will now allow you to remove visible watermarks from the images, videos, and music made with AI tools. With the update, you can toggle off a new "Media watermark" setting in Gemini and Google's AI video generator, Flow. When toggled off, Google will remove the "sparkle" watermark that appears in the bottom-right corner of […]
Turning off this setting won't affect invisible benchmarks used to identify an AI generated file.
Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer whi
Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose stream. Real-time generation is essential for interactive applications such as live streaming, telepresence, and virtual avatars, yet diffusion-based systems require minutes to hours per clip, precluding responsive interaction. We present LiveAnimate, to our knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT). A two-stage training pipeline first adapts a p
Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce
Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can ne
A new research paper frames AI adoption as a "tragedy of the cognitive commons." Every company that cuts entry-level jobs benefits individually, but the collective expertise of entire professions erodes. The consequences may not become visible until 2030 to 2045, when today's missing junior talent should have become tomorrow's experienced workforce. The article The "tragedy of the cognitive commons" explains how rational AI adoption could destroy entire professions' expertise appeared first on The Decoder .
Joi AI hired 10 people to masturbate using AI companions as part of a monthlong “wellness” study. The company claims the practice could help “solve male loneliness.”
We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs. We establish the recurrence of this organization across five linear attention archit
Rib fractures are common and time-consuming to localize on computed tomography (CT). We ask whether fractures detected independently in two orthogonal CT-derived projections (anteroposterior and lateral) can be paired across views and triangulated into reliable 3D points at a controlled rate of false outputs, and we answer it with a staged diagnostic study. The projection geometry is exact, and given correct correspondence, localization is accurate (median 4.0 mm, 88% within 10 mm, 93.6% rib-exact). On a sealed 55-case cohort, a large share of fractures is in principle recoverable (61.1% dual-
Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and costly expert trial and error. As pruning objectives, budgets, and model architectures diversify, manually navigating the expanding design space becomes increasingly difficult. This paper aims to build an AI4AI framework for visual-token pruning by addressing a natural question: Can large language models automatically design effective visual-token reduction algorithms? Although LLMs possess broad algorithmic knowled
When Twitch announced that streamers could opt out, thousands of users questioned why their content was being used to train AI models in the first place.
A plaintiff in Connecticut embedded invisible prompt injections in court filings, formatted in 3-point white text on a white background, to manipulate a potential AI review system. Judge Spader compared the attempt to secretly tampering with a jury and revoked the plaintiff's electronic filing privileges. The court stressed that Connecticut doesn't use AI to review filings, but the intent alone was enough to warrant sanctions. The article Plaintiff hid invisible AI instructions in court filings to secretly influence automated review appeared first on The Decoder .
World Labs, the startup founded by AI pioneer Fei-Fei Li, has unveiled a simulation engine that trains robot controllers entirely in virtual environments. From a single real-world task, the system generates thousands of controlled variations. The trained models then ran for one hour each on five different robot platforms without human intervention. How well the results hold up in more complex everyday situations remains to be seen. The article World Labs turns one real-world robot task into thousands of simulated variations for training appeared first on The Decoder .
The Unitree G1 has found online fame as a relatively affordable robot that can charm a crowd. But can it ever hold down a real job?
Instagram's wordmark is iconic. Well, was iconic. Apparently Instagram thought it looked old, so the company rolled out a new one this week. It doesn't look like the old Instagram wordmark. It doesn't even look like it spells Instagram anymore. And we cannot figure out why Instagram decided to do this. On this episode of […]
OpenAI's Computer History records clicks, keystrokes, and app switches on Mac and turns them into a searchable timeline for ChatGPT and Codex. The data is stored locally as unencrypted Markdown files. OpenAI says it's not used for AI training, but memories that feed into chats may still end up as training data. The article OpenAI's Computer History turns your clicks and keystrokes into a searchable ChatGPT memory timeline appeared first on The Decoder .
Here's the AI News you likely missed this week. Try Seedance 2.5 on Artlist here 👉 https://artlist.io/?artlist_aid=Mattwolfe_4110&utm_source=affiliate_p&utm_medium=Mattwolfe_4110&utm_campaign=Mattwolfe_4110 Discover More: 🛠️ Explore AI Tools & News: https://futuretools.io/ 📰 Weekly Newsletter: https://futuretools.io/newsletter Socials: ❌ Twiter/X: https://x.com/mreflow 🖼️ Instagram: https://instagram.com/mr.eflow 🧵 Threads: https://www.threads.net/@mr.eflow 🟦 LinkedIn: https://www.
Friend 2.0 might be the weirdest AI product launch I’ve seen all year 😭 The AI necklace can now talk back to you, and somehow the ad they made for it raises even more questions than the product itself. The founder has described Friend as a kind of “confidant, friend, God,” says the first 5,000 devices sold out, and now wants to sell 50,000 of these. And there are clearly people who genuinely want this. One commenter on X even said AI makes a better counselor than therapy. So… am I being old-
Natural gas prices could triple in some parts of the U.S., which could saddle hyperscalers with massive bills to power their AI data centers.
Anthropic is testing whether Claude Code can handle daily maintenance of the company's own apps, from crash fuzzing to dead-code removal. In a few weeks, the AI created 388 pull requests, and 46 percent were merged after human review. Claude Code inventor Boris Cherny sees this as "early signs of life that this might be possible." The article Claude Code now runs daily maintenance on Anthropic's software with a 46 percent merge rate appeared first on The Decoder .
Human-AI marriages are not currently recognized by US law. Some Republican state policymakers are drafting legislation to keep it that way.
Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day. Self-supervised foundation encoders change the regime: with a DINOv2 teacher, confidence saturates, so the filtering that helped a weak teacher can hurt a strong one. We propose CW-BASS v2, a saturation-aware pseudo-label selection method that reads the teacher's confidence regime rather than committing to one rule. It pa
We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm SE(3) transformations into attention via PRoPE-style geometric encoding, preserving arm identity and rigid-motion
Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step distillation accelerates generation by reducing denoising steps, while online control imposes a causal constraint: frames and blocks should depend on history and controls available during generation. Existing video distribution matching distillation (DMD) pipelines, however, often supervise causal few-step students using bidirectional teachers that score complete clips. The score for a target can therefore depend on future frames and controls that were unavailable when the student
Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operational, economic, and safety costs. Such events are rare in historical records, leaving insufficient training signal for machine learning models. Synthetic data augmentation offers a principled solution, but conventional generative models under-represent distributional tails and give no guarantee against operationally infeasible instances, such as a short air time paired with a long flight distance. No existing approach addresses both limitations for
Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a gendered target language such as Romanian. This bias results in translations that default to masculine forms or reinforce gender stereotypes. We propose a hybrid pipeline to mitigate this issue by combining large language model (LLM)-based gender classification with neural machine translation (NMT). Our system uses a fine-tuned LLM to detect the intended gender of target words in English sentences and insert inline gender hint tags. These tagged
No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for cons
We introduce , a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. consists of two coupled models: a prefiller Q, which leverages full attentionIn practice, we use interleaved full and sliding-window attention for Q, as this yields stronger performance. The essential requirement is that Q be more expressive than P, with access to the full history. to produce memory targets m'_t, and a decoder P, which uses only sliding-window attention and recurrent K/V injection to produce decoder memories m_t f
Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivations mislead users. We characterize this futile reasoning phenomenon through systematic analysis, revealing universal capability overreach and systematic miscalibration between capability and behavior. The dominant failure mode is specious reasoning, which outputs look superficially valid but contain subtle errors, escalating with task difficulty. To address this, we introduce CaRL (Capability-aligned Reinforcement Lea
Posts from the accounts I follow on X.
Nothing from this group today.
In my conversations with leaders at OpenAI, Anthropic, and Google, I'm just going to say that they don't seem to worry about context length/context windows.
— Allie K. Miller (@alliekmiller) 10h ago
Like, at all.
They all have crazy long-running threads.
One told me one of his threads has billions of tokens.
My Al Agent Workshop with @mcuban might be the best workshop we've ever had.
— Allie K. Miller (@alliekmiller) 13h ago
Over 20,000 people registered and my inbox is filled with people making REAL CHANGE in the way they work and what they're able to accomplish.
If you want a transformation that makes you DM "oh my god" https://t.co/iNrxtKIiJz
whats it called when the latest models make me think all previous versions are idiots?
— Ben Tossell (@bentossell) 14h ago
// Skill Misevolution in Self-Improving LLM Agents //
— DAIR AI (@dair_ai) 10h ago
Self-improving agents write their successes down as reusable skills. An unsafe success becomes reusable policy long after the input that triggered it is gone.
SkillMisevo-Gym versions skill state across agent frameworks so https://t.co/Fn2vNyOJkx
If you hand-tune agent harnesses, this one is worth your time.
— DAIR AI (@dair_ai) 14h ago
AutoDesign puts the harness itself inside the optimization loop. A meta-harness optimizer reads rollout feedback and directs a code agent to rewrite the harness, round after round.
They test it on paper-to-poster https://t.co/IEL39w77WU
I’ve been itching to do another livestream. Going to jump on live on Tuesday. What should we deep-dive on? What should we test out?
— Matt Wolfe (@mreflow) 14h ago
Another one. https://t.co/LU2ZOXPB7G
— Matt Wolfe (@mreflow) 14h ago
Apply or refer the best person you know: https://t.co/9xAZ2qKanM
— Rowan Cheung (@rowancheung) 14h ago
I'm HIRING 7 more people to help me build the future of AI media and education.
— Rowan Cheung (@rowancheung) 14h ago
The Rundown now has over 3M active readers, and our expansion plans far outpace what we can handle.
$2,000 referral bonus if you help us find someone we hire full-time.
Open roles and link to apply
A Beijing neurosurgeon resident just used GPT-5.6-Sol inside ChatGPT Work to prove a 22-year-old math conjecture.
— The Rundown AI (@TheRundownAI) 15h ago
Shanmu Jin was teaching himself the math for brain ultrasound research.
He locked the model off the web, started it, and left.
Sixteen hours later it had the https://t.co/E1stIRJLh4
Read more: https://t.co/d0XzMLcfjW
— The Rundown AI (@TheRundownAI) 15h ago
Top stories in tech today:
— The Rundown AI (@TheRundownAI) 15h ago
- Google wearables now track insulin resistance
- Flock tightens data rules after backlash
- Google makes Pixel a Gemini machine
- Joby buys its way into the defense tech race
- Quick hits on other tech news https://t.co/3yaTl7u7Pg
Read more: https://t.co/bBiflnnimq
— The Rundown AI (@TheRundownAI) 19h ago
Top stories in AI today:
— The Rundown AI (@TheRundownAI) 19h ago
- OpenAI previews frontier speed boost
- Rowan’s Corner: I asked AI to audit me
- Build a work ‘Second Brain’ that updates itself
- Anthropic's AI agents wage a turf war https://t.co/hL8ql4teKG
New: A map of the most important skills in AI Engineering. https://t.co/VVkn1Dqp1N
— Andrew Ng (@AndrewYNg) 13h ago
I was once in LA, and someone realized my last name was Valenzuela and asked me if I was related to Fernando Valenzuela.
— Cristóbal Valenzuela (@c_valenzuelab) 11h ago
Strangest question I’ve ever received. How would he know my dad’s name and ask me that? Why would you even know my dad’s name? Was this a joke? (He’s not
There was a knot so complex that no one had figured out how to untie it. The legend said that whoever managed to untie it would become the next emperor. Many had tried and all had failed.
— Cristóbal Valenzuela (@c_valenzuelab) 17h ago
Then Alexander the Great came across the knot. He looked at it, took out his sword, and cut https://t.co/ECAZcQ3dl2
Grok 4.6 runs The Gauntlet https://t.co/oDmQa9ah0P
— Elon Musk (@elonmusk) 1h ago
😂 https://t.co/3nQurlkWjQ
— Elon Musk (@elonmusk) 3h ago
Any censorship required by governments is now clearly visible https://t.co/eQMdoYlhkA
— Elon Musk (@elonmusk) 4h ago
Grok 4.6 now in Copilot https://t.co/kr6LQ9vyQS
— Elon Musk (@elonmusk) 4h ago
Orbital compute will be the only way to scale AI probably sometime in 2029 due to power availability& permitting problems on land https://t.co/nLTc6fuHXK
— Elon Musk (@elonmusk) 13h ago
Yes https://t.co/U2I4lY4RGN
— Elon Musk (@elonmusk) 15h ago
It is an honor to have such a great team join @SpaceX https://t.co/CyOaNE9ioy
— Elon Musk (@elonmusk) 15h ago
How to use @bot https://t.co/8myEssCRSL
— Elon Musk (@elonmusk) 1d ago
Something has happened with post-training as shown by DeepSeek flash & GLM-5.3 updates.
— Emad Mostaque (@EMostaque) 19h ago
Same base, big improvement in perf to frontier levels.
Can't explain this by even logit distillation etc
These are all hard benchmarks & GLM 5.3 is now top on by cyberdefense & GDPval! https://t.co/5Rfdjf1dnz
Suggestion for @X team
— Emad Mostaque (@EMostaque) 20h ago
If I have @grok Super Heavy let me use it to edit the code for my timeline directly up to certain parameters
I should be able to talk to @grok and have it present any way I want and look any way I want just from that conversation https://t.co/R3gMlPlRDS
try reservation search in chatgpt! https://t.co/RcBIeoht5V
— Greg Brockman (@gdb) 1h ago
hedonic treadmill of model expectations
— Greg Brockman (@gdb) 4h ago
wild to see Sol at 14x speed: https://t.co/9laJ2Xcvgw
— Greg Brockman (@gdb) 1d ago
chatgpt for saving you money https://t.co/rM7wGICiCi
— Greg Brockman (@gdb) 1d ago
Happy to see @NVIDIAAI released their expert models for MOPD. Starting to make research there much more accessible (tho these are big models for most researchers). https://t.co/LocLnse6HO
— Nathan Lambert (@natolambert) 7h ago
With prodding from @xeophon I added a crucial detail. Data industry go brrr in China.
— Nathan Lambert (@natolambert) 8h ago
Our interconnects group chat has on many occasions been discussing the data industry in China recently, a huge change from when we visited in April. https://t.co/bCaKPqZhYy https://t.co/UTPpVNdNEk
GLM 5.3 notes and why we should stop being so surprised about these very strong Chinese models (most of this is talking myself through some of my denial -- yes, these models are the real deal).
— Nathan Lambert (@natolambert) 9h ago
https://t.co/hePI3fILRF
Farewell Seattle! It's been such a wonderful life/career stage for me. Onto new adventures (and mountains). https://t.co/cVFr3ikyVY
— Nathan Lambert (@natolambert) 13h ago
Lots of people posting about Z ai “benchmaxxing” to make this model. I think the truth is messy and many faceted:
— Nathan Lambert (@natolambert) 15h ago
1. Yes Zai probably cares slightly more about public benchmarks than OpenAI/Ant, helps with marketing
2. Zai is not benchmaxxing to the point where the model is https://t.co/esFrQ11EJq
POV: You’re romanticizing a New York City lifestyle. 🗽🎞️
— Adobe Firefly (@AdobeFirefly) 8h ago
Try the full prompt pack:
Top Right
https://t.co/Dx0zWLp8P2
Bottom Right
https://t.co/FNuBd8x84M
Bottom Left
https://t.co/b946h7syBI https://t.co/K34qqt1NzG
We’ve written an FAQ to answer some of the questions we've received about watermarking.
— Anthropic (@AnthropicAI) 11h ago
In summary:
• We’re implementing watermarking to comply with the EU AI Act. Other major model developers have signed the same Code of Practice and will also be implementing watermarking;
As part of our Responsible Scaling Policy, we publish regular Risk Reports. These share detailed information on the risks of our systems and how prepared we are to address them.
— Anthropic (@AnthropicAI) 12h ago
Our second Risk Report is now available: https://t.co/NgWnDmXZD3
Learn more: https://t.co/O0OtPGuYHj
— NVIDIA AI (@NVIDIAAI) 11h ago
ICYMI: Not every step in an agent workflow needs the same model.
— NVIDIA AI (@NVIDIAAI) 11h ago
Meet NVIDIA NeMo Switchyard, a new open-source library for model routing.
Use frontier models for complex reasoning and planning, and NVIDIA Nemotron Lightning for high-volume, specialized execution. https://t.co/TI3dCwDrfS
We built Pika Audio Models around a simple engineering challenge: make frontier audio quality more accessible at scale.
— Pika (@pika_labs) 7h ago
Our team's efficient inference techniques made Soundtrack, Music, SFX, and Speech up to 20× more cost-efficient than other audio models.
Read the technical https://t.co/rgBGrq3ocE
Try all 4 of our Audio Models through the Pika API Club https://t.co/TdCmKVtKBo
— Pika (@pika_labs) 7h ago
We’re thrilled to be able offer prices this low, thanks to our team’s innovations in training and inference efficiency. A few highlights:
— Pika (@pika_labs) 10h ago
• Pika Soundtrack is 0.617 / seconds and 2x more cost-efficient than Hunyuan Foley, the only model with comparable video-to-audio
The Pika Audio model family is just one of the things we’ve been working on behind the scenes. Keep an eye on us for more ways we’ll be making gen media more accessible to more people…
— Pika (@pika_labs) 10h ago
ICYMI https://t.co/TdCmKVtKBo
— Pika (@pika_labs) 10h ago
Get started: https://t.co/j1NtDVkhuu
— Runway (@runwayml) 11m ago
Seedance 2.5 in 1080p is now live on Runway. Early access starts today, bringing sharper detail at higher resolution.
— Runway (@runwayml) 11m ago
Get started at the link below. https://t.co/S2tNz0XKEZ
Watch all the winning films at: https://t.co/1lOfrKxKKn
— Runway (@runwayml) 10h ago
Fifth Place
— Runway (@runwayml) 10h ago
Eau de Paw — XAZINGA https://t.co/6XNI96gviL
Fourth Place
— Runway (@runwayml) 10h ago
Solace x Vlad — @laszlogaal_ https://t.co/S1Vx1msii3
What ran today and what it cost.