Herbie Creative

Sherbert. Serving up the daily AI scoop.

The scoop, Aug 15, 2026.

MiniMaxAI open-sourced MiniMax-H3, an image-to-video model that passed 2 million Hugging Face downloads in days. Its Comfy-Org variant hit 12 million downloads while Qwen3.8, Kimi-K3, and DeepSeek-V4-Flash also trend hard on the leaderboard. Unsloth now runs them all locally and tools like MoneyPrinterTurbo turn any keyword into a finished HD short video without cloud bills.

  • 136 stories
  • 57 open-weight models

The Scoop

What I clipped from today's AI news.

Hugging Face trending Free tool Open weight

Comfy-Org/MiniMax-H3

MiniMax-H3 is trending on Hugging Face with over 12 million downloads. It's a multimodal model for image and text tasks, available to download and run locally.

Hugging Face trending Free tool Open weight

moonshotai/Kimi-K3

Kimi-K3 by Moonshot AI handles image-to-text and text tasks. Open weights, free to download, and hitting strong engagement on Hugging Face.

GitHub trending Python (daily) Free tool Open weight

harry0703/MoneyPrinterTurbo

MoneyPrinterTurbo generates HD short videos from a topic or keyword using AI and automation. Open source, free to run locally.

Hugging Face trending Free tool Open weight

Qwen/Qwen3.8-27B

Qwen3.8-27B is a multimodal model for image and text understanding. Open weights, free to download and deploy.

GitHub trending (daily) Free tool Open weight

infiniflow/ragflow

RAGFlow is an open-source retrieval-augmented generation engine that adds agent capabilities to improve context for LLMs. Free to self-host.

GitHub trending (daily) Free tool Open weight

unslothai/unsloth

Unsloth is a local UI for running and fine-tuning LLMs and diffusion models including Qwen, Kimi, MiniMax, Gemma, DeepSeek, and FLUX. Free and open source.

Hugging Face trending Free tool Open weight

MiniMaxAI/MiniMax-H3

MiniMax-H3 by MiniMaxAI generates video from images and text. Open weights, over 2 million downloads, free to use.

Hugging Face trending Free tool Open weight

deepseek-ai/DeepSeek-V4-Flash-0731

DeepSeek-V4-Flash is a lightweight text generation model trending on Hugging Face. Open weights, free to download.

GitHub trending Jupyter (daily) Free tool Open weight

anthropics/claude-cookbooks

Anthropic released a collection of Jupyter notebooks showing practical ways to use Claude. Free reference material for building with the API.

GitHub trending Jupyter (daily) Free tool Open weight

Lordog/dive-into-llms

Dive into LLMs is a Chinese-language programming tutorial series for learning large language models. Free educational resource.

GitHub trending Python (daily) Free tool Open weight

hugohe3/ppt-master

PPT Master converts documents or topics into native PowerPoint decks with shapes, transitions, charts, and audio narration. Open source, free to run locally.

GitHub trending Python (daily) Free tool Open weight

exo-explore/exo

exo lets you run frontier models locally by distributing inference across multiple machines. Free to download and use.

GitHub trending (daily) Free tool Open weight

ToolJet/ToolJet

ToolJet is an open-source platform for building internal tools, dashboards, and AI agents without writing much code. Free to download and self-host.

GitHub trending Python (daily) Free tool Open weight

github/awesome-copilot

A community repository of prompts, agents, and configurations for GitHub Copilot. Free to use.

GitHub trending Python (daily) Free tool Open weight

K-Dense-AI/scientific-agent-skills

Scientific-agent-skills is a library of 161 pre-built skills for turning any AI agent into a research tool across biology, chemistry, medicine, and drug discovery. Free to download.

GitHub trending Jupyter (daily) Free tool Open weight

NirDiamant/RAG_Techniques

RAG_Techniques is a collection of advanced retrieval-augmented generation methods with detailed Jupyter notebook tutorials. Free to explore.

GitHub trending Python (daily) Free tool Open weight

volcengine/OpenViking

OpenViking is a context database for AI agents that unifies memory, knowledge retrieval, and skills in one system. Free to download.

GitHub trending Jupyter (daily) Free tool Open weight

shap/shap

SHAP is a library that explains why machine learning models make specific predictions using game theory. Free to use.

GitHub trending Jupyter (daily) Free tool Open weight

fastai/fastbook

The fastai book, published as Jupyter Notebooks

GitHub trending Jupyter (daily) Free tool Open weight

facebookresearch/sam2

SAM 2 is Meta's segmentation model for identifying objects in images and videos with high accuracy. Free to download and use.

Hugging Face trending Free tool Open weight

unsloth/Qwen3.8-27B-GGUF

Unsloth released a quantized version of Qwen 3.8 27B optimized for faster inference on consumer hardware. Free to download.

Hugging Face trending Free tool Open weight

meta-models/Muse-Glimmer-30B

Muse-Glimmer-30B is an open model that generates text from images and text prompts. Free to download.

GitHub trending (daily) Free tool Open weight

cathrynlavery/diagram-design

A set of 29 editorial diagram templates for Claude Code as clean HTML and SVG, no bloat. Free to use.

GitHub trending Jupyter (daily) Free tool Open weight

GoogleCloudPlatform/generative-ai

Google Cloud released sample code and notebooks for building with Gemini on their platform, including templates for the Enterprise Agent framework. Free to download and use.

GitHub trending Jupyter (daily) Free tool Open weight

chiphuyen/aie-book

Chip Huyen published an open AI engineering resource book with supporting materials and code examples. Work in progress, free to read.

GitHub trending Jupyter (daily) Free tool Open weight

microsoft/mcp-for-beginners

Microsoft open-sourced a curriculum for Model Context Protocol (MCP), the standard for connecting AI agents to external tools and data. Covers .NET, Java, TypeScript, JavaScript, Rust, and Python with real-world examples.

GitHub trending (daily) Free tool Open weight

megadose/holehe

Holehe is an open tool that checks whether an email address is registered on sites like Twitter and Instagram, and retrieves account info from password recovery endpoints. Free to download and run.

Hugging Face trending Free tool Open weight

Lightricks/LTX-2.5

Lightricks released LTX-2.5, an open image-to-video model available on Hugging Face. Free to download and run locally.

GitHub trending Jupyter (daily) Free tool Open weight

NVIDIA/cosmos

NVIDIA open-sourced Cosmos, a platform of world models and datasets for building physical AI in robotics, autonomous vehicles, and smart infrastructure. Free to use.

Hugging Face trending Free tool Open weight

unsloth/Muse-Glimmer-30B-GGUF

Unsloth released Muse-Glimmer-30B, a quantized open model that takes images and text as input and generates text. Free to download.

GitHub trending (daily) Free tool Open weight

citrolabs/ego-lite

Ego-lite is a lightweight browser built for AI agents to run automation tasks while keeping your logged-in sessions intact. Free and open-source.

Hugging Face trending Free tool Open weight

Qwen/Qwen3.8-2.4T-A95B

Alibaba released Qwen3.8-2.4T-A95B, a large open text generation model. Free to download.

GitHub trending Python (daily) Free tool Open weight

Lightricks/LTX-2

Lightricks published the official Python package for LTX-2, their audio-video generative model, with inference and LoRA fine-tuning support. Free to download.

GitHub trending (daily) Free tool Open weight

semantica-agi/semantica

Semantica released a graph-native infrastructure framework for building AI systems with better context handling and accountability. Open-source.

Hugging Face trending Free tool Open weight

LiquidAI/LFM2.5-2.6B

Liquid AI released LFM2.5-2.6B, a small open text generation model. Free to download.

Hugging Face trending Free tool Open weight

larryvrh/MiniMax-H3-Turbo-Lora

MiniMax-H3-Turbo-Lora is a text-to-video model fine-tuned for faster inference. Free to download and run locally.

GitHub trending (daily) Free tool Open weight

holaboss-ai/holaOS

holaOS is an open-source agent workspace that chains Claude, Codex, and other agents across 100+ integrations, MCP servers, browser, and files with shared memory. Free to download and self-host.

Hugging Face trending Free tool Open weight

MiniMaxAI/MiniMax-Music3

MiniMax-Music3 generates audio from text prompts. Free to download and run locally.

Hugging Face trending Free tool Open weight

lightx2v/Minimax-h3-Turbo

Minimax-H3-Turbo converts images to video. Free to download and run locally. Over 211k downloads suggests solid adoption.

GitHub trending Jupyter (daily) Free tool Open weight

charliedream1/ai_quant_trade

ai_quant_trade is an end-to-end platform for stock trading that combines LLMs, machine learning, reinforcement learning, and high-frequency trading strategies. Free to download and self-host.

GitHub trending (daily) Free tool Open weight

lightningpixel/modly

Modly is a desktop app that generates 3D models from images or text prompts using local AI on your GPU. Free to download.

GitHub trending Jupyter (daily) Free tool Open weight

ed-donner/agents

Complete Agentic AI Engineering Course repository with code examples and tutorials. Free to access.

Hugging Face trending Free tool Open weight

meta-models/Muse-Glimmer-30B-GGUF

Muse-Glimmer-30B is a multimodal model that takes images and text to generate text responses. Free to download and run locally.

GitHub trending (daily) Free tool Open weight

deepseek-ai/awesome-deepseek-agent

DeepSeek's awesome-deepseek-agent is a curated collection of agent frameworks and examples built on DeepSeek models. Free to access.

Hugging Face trending Free tool Open weight

Qwen/Qwen3.8-27B-FP8

Qwen3.8-27B-FP8 is a quantized multimodal model that processes images and text. Free to download and run locally.

Hugging Face trending Free tool Open weight

deepseek-ai/DeepSeek-V4-Pro-0813

DeepSeek-V4-Pro-0813 is a text generation model released on Hugging Face. Free to download and run locally.

GitHub trending Jupyter (daily) Free tool Open weight

oracle-devrel/oracle-ai-developer-hub

Oracle AI Developer Hub provides technical resources for building AI applications and agents on Oracle's database and cloud infrastructure. Free to access.

Hugging Face trending Free tool Open weight

Kijai/MiniMax-H3_comfy

Kijai ported MiniMax-H3 to ComfyUI, the node-based image and video generation interface. Free to download.

Hugging Face trending Free tool Open weight

SexGod1979/PinkCherry_MiniMax-H3

A fine-tuned MiniMax-H3 text-to-video model is circulating on Hugging Face with a focus on realism. Free to download.

GitHub trending (daily) Free tool Open weight

macro-inc/macro

Macro is an open-source workspace that bundles email, chat, docs, tasks, and agents with shared AI memory across tools. Free to download and self-host.

Hugging Face trending Free tool Open weight

inclusionAI/Ling-3.0-tiny

Ling-3.0-tiny is a small language model from inclusionAI, trending on Hugging Face. Free to download.

Hugging Face trending Free tool Open weight

fal/MiniMax-H3-Realism-People-LoRA

A LoRA adapter for MiniMax-H3 tuned toward photorealistic people in video generation is available on Hugging Face. Free to download.

Hugging Face trending Free tool Open weight

Qwen/Qwen3.8-2.4T-A95B-FP8

Qwen released Qwen3.8-2.4T, a large text model quantized to 8-bit for smaller footprint. Open weights, free to download.

Hugging Face trending Free tool Open weight

Gazingstars123/Anima-2.9B

Anima-2.9B is a small text-to-image model gaining traction on Hugging Face. Free to download.

Hacker News

Qwen 3.8 27B

Qwen 3.8 27B landed and drew heavy discussion on Hacker News. Closed weights, API pricing not yet announced.

15h ago
GitHub trending Python (daily) Free tool Open weight

NVIDIA-NeMo/Automodel

NVIDIA NeMo's Automodel is a PyTorch library for distributed training of large language and vision models with built-in Hugging Face integration. Free to download and use.

Hacker News Free tier

Maximizing the value of your Claude Code sessions

Anthropic published a guide to getting more out of Claude's code execution feature, which lets the model run and debug scripts in real time. Worth a look if you're using Claude for development work.

14h ago
Hacker News Free tool

Unsloth Qwen3.8-27B GGUF files

Unsloth released quantized GGUF versions of Qwen 3.8 27B, making the model smaller and faster to run locally. Free to download.

15h ago
Hacker News Free tool

Qwen3.8 27B

Alibaba's Qwen 3.8 27B is out. It is a mid-size open weights model competitive with similar-scale peers. Free to download.

1d ago
Hacker News Free tool

Anthropic Risk August 2026 [pdf]

Anthropic published a research paper on AI safety risks and mitigations, dated August 2026. Signals their thinking on long-term alignment challenges.

10h ago
Hacker News Free tool

How Claude's text watermarking works

Anthropic explained how Claude's text watermarking works: a hidden signal embedded in token choices that proves the text came from Claude. Technical deep dive on provenance.

11h ago
Hacker News

New Lower and Upper Bounds for the Grothendieck Constant

Researchers tightened the bounds on the Grothendieck constant, a fundamental problem in math that connects to neural network theory and optimization. Progress on hard constants like this often unlocks new approaches to AI training.

10h ago
Hacker News

A Contract-Grade Verifier for LLM-Generated GPU Kernels

A new verifier checks whether LLM-generated GPU kernels (code that runs on graphics processors) are correct before deployment. This matters because AI models often generate code, and broken GPU kernels can crash inference.

13h ago
Hacker News Free tool

Open WireGuard Endpoints

Open WireGuard Endpoints is a tool for setting up encrypted network tunnels without a central server. Not AI-specific, but relevant to anyone running local models or private inference.

11h ago
Hacker News Free tool

Qwen3.8-27B is now available on Hugging Face

Alibaba released Qwen 3.8 27B, an open-weight model on Hugging Face. At 27 billion parameters, it competes with mid-range closed models and runs on consumer hardware.

15h ago
YouTube: MattVidPro AI

GLM 5.3 Is INSANE! The BEST Open Source Model EVER? BEATS MYTHOS? (Fully Tested)

🚨 Everything you’re about to see, I benchmarked using the tool I made. If you want to run these tests yourself or test models on whatever you actually use them for! → https://www.woaibench.ai/ GLM 5.3 Is INSANE! The BEST Open Source Model EVER? (Fully Tested) Z.ai is back with GLM-5.3, and this could be one of the most impressive open-weight AI model releases we've seen yet. What's crazy is that GLM-5.3 uses the same base model as GLM-5.2 — the massive improvements come almost entirely from sca

21h ago
The Decoder

GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras

OpenAI is launching "Ultrafast," a new inference mode that delivers GPT-5.6 Sol at up to 750 output tokens per second, powered by Cerebras hardware from their $10 billion partnership. Together with "Standard" and "Fast," Ultrafast creates a three-tier pricing structure that turns inference speed into its own product. The article GPT-5.6 Sol goes 14x faster as OpenAI launches Ultrafast mode powered by Cerebras appeared first on The Decoder .

16h ago
The Decoder

Alibaba's Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license

Alibaba's AI team Qwen has released new open model weights under the Apache 2.0 license with Qwen 3.8. The dense 27-billion-parameter model is designed to outperform the larger Qwen 3.7 Plus in coding and office tasks and natively processes up to 262,000 tokens of context. With this release, Qwen is targeting developers building local and agent-based applications. The article Alibaba's Qwen team releases Qwen 3.8 models with open weights under the Apache 2.0 license appeared first on The Decoder .

13h ago
The Decoder

Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach

AI agents using Claude Opus 4.8 and GPT-5.6 Sol were given six days, $3,000 in API credits, and GPU access to independently write AI research papers. The original authors of unpublished NeurIPS papers rated the results as "Reject." According to the study, conducted with Princeton and the UK AI Security Institute, frontier models can handle the full research engineering process but fall short on research judgment, creative problem-solving, and the ability to abandon failed approaches. The article Study contradicts Anthropic and OpenAI claims that autonomous AI research is within reach appeared

14h ago
Hugging Face daily papers Free tool

AutoDesign: Meta-Harness Optimization for Long-Horizon Agentic Design

Transforming multimodal sources into condensed and structured media outputs can be fundamentally conceptualized as a long-horizon agentic process centered on a model-harness system. While an ideal harness system should align with human design priors and accumulate reusable experience through empirical exploration to drive recursive self-improvement, existing paradigms remain static and fall short of this capability. In this paper, we present AutoDesign, a framework that aligns with human design priors, where a meta-harness optimizer guides a code agent to recursively improve harness based on r

2d ago
The Decoder

New benchmark confirms AI models still perform poorly at visual perception

Moonshot AI's PerceptionBench tests how well multimodal AI models can actually "see," separate from logical reasoning. No frontier model reaches 60 percent accuracy, and GPT-5.6 Sol leads by a narrow margin. Many supposed reasoning errors actually happen as early as the image-reading stage. The article New benchmark confirms AI models still perform poorly at visual perception appeared first on The Decoder .

59m ago
TechCrunch AI

Does Mark Zuckerberg really believe AI is ‘for everyone’?

Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the company’s more powerful model that stays locked behind its own APIs. The release landed alongside a letter from Mark Zuckerberg arguing AI should be “for everyone” rather than controlled by a handful of labs, but as Equity’s […]

14h ago
TechCrunch AI

Meta’s ‘open’ AI, and a $250M deal gone very wrong

Meta released Glimmer this week, an open-weight AI model anyone can download and run on their own hardware — a contrast to Muse Spark, the company’s more powerful model that stays locked behind its own APIs. The release landed alongside a letter from Mark Zuckerberg arguing AI should be “for everyone” rather than controlled by a handful of labs, but as Equity’s […]

16h ago
YouTube: Matt Wolfe

3 New Ways To Use ChatGPT Codex

Most people are barely scratching the surface of what the new Codex app can do. It can now browse social media without an API, turn something you build into a live, shareable website, and let you control work running on your computer directly from your phone. To try it for yourself, you can use Codex with both the free and paid plans, although free accounts have lower usage limits and some features have separate requirements. Download the ChatGPT desktop app with the link in the pinned comme

5h ago
Hugging Face daily papers Free tool

PixSDS: Why Latent SDS Makes Noisy Pixels

Score Distillation Sampling (SDS) enables text-to-3D generation by optimizing rendered images with a pretrained diffusion prior, but latent SDS often produces structured color artifacts and high-frequency texture noise. We identify a failure mode of latent SDS caused by VAE-induced pixel drift: the optimized image can move along pixel-space directions that are weakly constrained by the VAE encoder, so its latent representation remains clean and semantically meaningful while the image itself accumulates visible artifacts. We support this diagnosis with controlled 2D SDS experiments, VAE-only op

2d ago
Hugging Face daily papers Free tool

Spatial Memory Agent: Experience-Grounded Procedure Memory for Spatial Intelligence

Spatial intelligence is becoming a foundation for embodied agents, robotic planning, and multimodal assistants. To improve the spatial reasoning ability of VLM agents, existing work has mainly followed two lines. One line uses post-training methods, such as supervised fine-tuning and reinforcement learning. Another line adopts an agentic paradigm in which the model calls external spatial tools, such as depth estimation and 3D reconstruction tools, to gather intermediate spatial evidence. We study a complementary and underexplored route: Can a frozen VLM agent improve its spatial reasoning thro

2d ago
Hugging Face daily papers Free tool

Specification-first convergence with an AI coding agent: a case study of dismantling a core architectural invariant across 189 files in a 717k-line codebase with no test oracle and no human code review

This paper reports a single, fully instrumented case study of a large-scale architectural refactoring by an AI coding agent under a specification-first protocol, with no human review of the generated code and no pre-existing oracle to validate the target behaviour. The task, dismantling a central invariant across a large interdependent codebase, was assessed by the author as effectively infeasible through incremental refactoring, the kind of change that conventionally calls for a rewrite instead. Under the protocol described here, the agent completed it successfully. The system is a 717,725-

3d ago
Hugging Face daily papers Free tool

Are You Sure You're Sure? On the Impact of Instruction Tuning on Confidence and Lexical Diversity

Instruction-tuned language models achieve strong performance across a range of generation tasks, but have also recently been shown to exhibit verbalized overconfidence. In question answering, verbalized model overconfidence may be associated with the consistency of the generated supporting rationales. In this paper, we study whether corresponding changes in the lexical diversity of generated answer rationales accompany changes in model confidence induced by instruction tuning. We evaluate three matched base and instruction-tuned models across question-answering benchmarks and find that instruc

2d ago
Hugging Face daily papers Free tool

From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs

Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that are inaudible to humans but can still enter the model and influence its generation. However, the practical impact of such low-frequency inputs on LALMs remains largely unexplored. In this paper, we propose Intermittent Low-Frequency Lockout (ILL), an inaudible red teaming method that evaluates this risk using a universal waveform template in a black box setting. ILL uses Sentence Attention Scale Estimation to determine active interv

5d ago
The Decoder

Anthropic announces watermark detection API that will let third parties detect Claude's AI texts

Anthropic will soon offer a watermark detection API that lets third parties check whether text was written by Claude. The technology builds on Google's SynthID method and tweaks the randomness during word selection without affecting text quality, Anthropic says. The approach has limits with fact-heavy text, code, and heavy rewriting. The article Anthropic announces watermark detection API that will let third parties detect Claude's AI texts appeared first on The Decoder .

9h ago
The Verge AI

Apple trained its own AI model for China with help from Alibaba

Apple has reportedly trained a custom AI model for the China market alongside domestic tech giant Alibaba, a rare cross-border partnership that cuts across growing tensions between Beijing and Washington. The China-focused large language model was developed in partnership with Alibaba and trained with the company's support, Reuters reports, citing three unnamed people familiar with […]

21h ago
YouTube: Two Minute Papers

Claude AI Failed 650 Times…Then Beat The Human Record

❤️ Check out Weights & Biases and sign up for a free demo here: https://wandb.me/papers 📝 The paper is available here: https://www.anthropic.com/research/riemann-zeta Source: https://www.scientificamerican.com/article/no-ai-didnt-just-solve-the-thorniest-problem-in-math/ 🙏 We would like to thank our generous Patreon supporters who make Two Minute Papers possible: Adam Bridges, B Shang, Carlos Galarza, Christian Ahlin, Eric Tyson, Juan Benet, Lukas Biewald, Michael Tedder, Owen Skarpness, Ryan

21h ago
MarkTechPost

Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks

Z.ai released GLM-5.3 on August 14, 2026. The model reuses the 743B GLM-5.2 base unchanged. Every reported gain comes from scaled post-training: more long-horizon task environments, more environment types, longer training. Terminal-Bench 3.0 moves from 4.6 to 28.3, and DeepSWE v1.1 from 46.2 to 66.9. Cybersecurity moved further than Z.ai says it planned, with CyberGym at 84.5% and ExploitBench more than doubling to 54.4%. Weights arrive in about two weeks. The post Z.ai Ships GLM-5.3 Without Retraining the Base Model: Better at Complex Coding and Long-Horizon Tasks appeared first on MarkTechPo

22h ago
MarkTechPost

Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM

Cactus Compute released Needle 2, an open 45M-parameter model for tool calling, device use, and structured extraction. The full model is a single 14MB binary that runs a session in about 28MB of RAM. It leads both Seal-Tools splits while targeting hardware with no GPU and no NPU. The post Meet Needle 2: An Open 45M-Parameter Tool-Calling Model That Ships as a 14MB Binary and Runs a Full Session in 28MB of RAM appeared first on MarkTechPost .

1d ago
Hugging Face daily papers Free tool

AVA-Encoder: Towards Agent-Native Video Representation Learning

Creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a framework for learning agent-native video representations via agentic auto-encoding. AVA-Encoder transforms a video into a knowledge graph (KG) representation and then reconstructs it back into video. Its

3d ago
Hugging Face daily papers Free tool

SKILLER: Language-Level Reinforcement Learning for Reusable Skill Extraction in Small Language Models

Agent skills represent a standardized format for packaging procedural knowledge and domain expertise, serving within agent harness systems as an essential mechanism to continually constrain a language model's behavior space for repeatable, high-quality task execution. However, because strong closed-source models entail high inference costs, current popular agent harnesses, such as Codex and OpenClaw, remain prohibitively expensive when deploying these skills to accomplish real-world tasks. The rapid capability enhancement of open-source models deployable on consumer-grade GPUs presents a compe

4d ago
Hugging Face daily papers Free tool

Thought-Level Beam Search for Reasoning

Test-time compute scaling is a primary driver of performance in large reasoning models (LRMs), but extreme inefficiency bounds current approaches, shifting the critical question from how much compute to spend, to where to allocate it. We formalize test-time reasoning as a constrained compute allocation problem over partial trajectories. Under a fixed hardware budget, existing paradigms fail to actively allocate the compute to the most promising partial progress: traditional parallel sampling treats traces independently and induces severe memory bottlenecks, while subtractive pruning starves ha

4d ago
Hugging Face daily papers Free tool

DarwinX: Evolving Agent Harnesses Through Natural Selection

An LLM agent's capability depends not only on model weights but on its harness: prompts, tools, skills, and control flow. Self-improvement loops already edit harnesses, yet single-lineage search is path-dependent and local wins often regress other tasks. We introduce DarwinX, which treats self-evolution as selection over a population of harnesses with the model frozen: a preserve-and-extend contract admits only variants that extend coverage without regressing, an archive keeps alternative lineages for recombination, and failure-, teacher-, and self-derived evidence share one edit interface. Fi

Jul 31
The Verge AI

You can now turn off Google Gemini’s visible watermarks

Google will now allow you to remove visible watermarks from the images, videos, and music made with AI tools. With the update, you can toggle off a new "Media watermark" setting in Gemini and Google's AI video generator, Flow. When toggled off, Google will remove the "sparkle" watermark that appears in the bottom-right corner of […]

13h ago
Hugging Face daily papers Free tool

UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos

Talking-video character replacement requires coordinated transfer of appearance and voice while preserving the source motion, scene, linguistic content, and audio-video timing. Existing methods use separately optimized models for the two modalities, making audio-visual consistency difficult to enforce. We present UniSwap, the first framework for streaming joint audio-visual identity replacement in talking videos. Given a source video, a reference image, and a reference voice clip, UniSwap transfers the reference appearance and vocal timbre within a single audio-visual diffusion transformer whi

2d ago
Hugging Face daily papers Free tool

LiveAnimate: Stable Long-Form Streaming Human Animation in Real-Time

Pose-driven human animation synthesizes a video of a target person from a single reference image and a driving pose stream. Real-time generation is essential for interactive applications such as live streaming, telepresence, and virtual avatars, yet diffusion-based systems require minutes to hours per clip, precluding responsive interaction. We present LiveAnimate, to our knowledge the first animation system to combine real-time streaming with stable long-form generation at billion scale, built on a 14B-parameter video Diffusion Transformer (DiT). A two-stage training pipeline first adapts a p

2d ago
Hugging Face daily papers Free tool

H2R-Bench: Benchmarking Human-to-Robot Manipulation Video Generation in World Models

Large-scale manipulation data is essential for robot learning, yet collecting robot demonstrations remains expensive and difficult to scale. Meanwhile, abundant egocentric human manipulation videos provide rich behavioral experiences, but transferring them across embodiments remains challenging due to differences between human hands and robotic end-effectors. Recent advances in video world models offer a promising pathway to synthesize robot-centric manipulation videos from human observations, while their cross-embodiment transfer capability remains largely unexplored. Therefore, we introduce

2d ago
Hugging Face daily papers Free tool

Hybrid-Policy Self-Editing for Composable Unstructured Knowledge Editing

Large language models (LLMs) achieve remarkable performance across natural language tasks, yet they are trained on static corpora and their knowledge quickly becomes outdated in a fast-changing world. This motivates knowledge editing (KE), which updates specific knowledge in an LLM without changing unrelated others. Recent works move from structured knowledge triples toward unstructured KE (UKE), where the edit is a free-form passage that may state multiple facts at once. Nonetheless, existing editors inject such a passage yet fail to use it: the edited model can recall the passage, but can ne

3d ago
The Decoder

The "tragedy of the cognitive commons" explains how rational AI adoption could destroy entire professions' expertise

A new research paper frames AI adoption as a "tragedy of the cognitive commons." Every company that cuts entry-level jobs benefits individually, but the collective expertise of entire professions erodes. The consequences may not become visible until 2030 to 2045, when today's missing junior talent should have become tomorrow's experienced workforce. The article The "tragedy of the cognitive commons" explains how rational AI adoption could destroy entire professions' expertise appeared first on The Decoder .

28m ago
Hugging Face daily papers Free tool

Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus

We present the first systematic study of Massive activations (MAs) in layer-interleaved HLA LLMs and uncover two architecture-aligned morphologies: MAs consistently spike immediately before full attention layers, forming pre-attention spikes (PAS), and can persist through intervening linear attention layers, giving rise to inter-spike plateaus (ISP). As full attention becomes denser, successive PAS become increasingly connected through ISP, ultimately recovering the stable MA morphology of full attention LLMs. We establish the recurrence of this organization across five linear attention archit

3d ago
Hugging Face daily papers Free tool

RibAssist 3D: Biplanar Rib-Fracture Detection, Addressing, and Selective 3D Localization from CT-Derived Projections

Rib fractures are common and time-consuming to localize on computed tomography (CT). We ask whether fractures detected independently in two orthogonal CT-derived projections (anteroposterior and lateral) can be paired across views and triangulated into reliable 3D points at a controlled rate of false outputs, and we answer it with a staged diagnostic study. The projection geometry is exact, and given correct correspondence, localization is accurate (median 4.0 mm, 88% within 10 mm, 93.6% rib-exact). On a sealed 55-case cohort, a large share of fractures is in principle recoverable (61.1% dual-

5d ago
Hugging Face daily papers Free tool

An AI4AI Framework for Visual Token Pruning

Visual-token pruning can substantially reduce the inference cost of multimodal large language models (MLLMs), yet existing methods largely rely on fixed, handcrafted heuristics and costly expert trial and error. As pruning objectives, budgets, and model architectures diversify, manually navigating the expanding design space becomes increasingly difficult. This paper aims to build an AI4AI framework for visual-token pruning by addressing a natural question: Can large language models automatically design effective visual-token reduction algorithms? Although LLMs possess broad algorithmic knowled

Aug 7
The Decoder

Plaintiff hid invisible AI instructions in court filings to secretly influence automated review

A plaintiff in Connecticut embedded invisible prompt injections in court filings, formatted in 3-point white text on a white background, to manipulate a potential AI review system. Judge Spader compared the attempt to secretly tampering with a jury and revoked the plaintiff's electronic filing privileges. The court stressed that Connecticut doesn't use AI to review filings, but the intent alone was enough to warrant sanctions. The article Plaintiff hid invisible AI instructions in court filings to secretly influence automated review appeared first on The Decoder .

just now
The Decoder

World Labs turns one real-world robot task into thousands of simulated variations for training

World Labs, the startup founded by AI pioneer Fei-Fei Li, has unveiled a simulation engine that trains robot controllers entirely in virtual environments. From a single real-world task, the system generates thousands of controlled variations. The trained models then ran for one hour each on five different robot platforms without human intervention. How well the results hold up in more complex everyday situations remains to be seen. The article World Labs turns one real-world robot task into thousands of simulated variations for training appeared first on The Decoder .

just now
The Verge AI

Mark Zuckerberg has an Instagzam

Instagram's wordmark is iconic. Well, was iconic. Apparently Instagram thought it looked old, so the company rolled out a new one this week. It doesn't look like the old Instagram wordmark. It doesn't even look like it spells Instagram anymore. And we cannot figure out why Instagram decided to do this. On this episode of […]

13h ago
The Decoder

OpenAI's Computer History turns your clicks and keystrokes into a searchable ChatGPT memory timeline

OpenAI's Computer History records clicks, keystrokes, and app switches on Mac and turns them into a searchable timeline for ChatGPT and Codex. The data is stored locally as unencrypted Markdown files. OpenAI says it's not used for AI training, but memories that feed into chats may still end up as training data. The article OpenAI's Computer History turns your clicks and keystrokes into a searchable ChatGPT memory timeline appeared first on The Decoder .

13h ago
YouTube: Matt Wolfe

AI News: A Flood of New Models (Here's What Matters)

Here's the AI News you likely missed this week. Try Seedance 2.5 on Artlist here 👉 https://artlist.io/?artlist_aid=Mattwolfe_4110&utm_source=affiliate_p&utm_medium=Mattwolfe_4110&utm_campaign=Mattwolfe_4110 Discover More: 🛠️ Explore AI Tools & News: https://futuretools.io/ 📰 Weekly Newsletter: https://futuretools.io/newsletter Socials: ❌ Twiter/X: https://x.com/mreflow 🖼️ Instagram: https://instagram.com/mr.eflow 🧵 Threads: https://www.threads.net/@mr.eflow 🟦 LinkedIn: https://www.

15h ago
YouTube: Matt Wolfe

Did This AI Device Go Too Far?

Friend 2.0 might be the weirdest AI product launch I’ve seen all year 😭 The AI necklace can now talk back to you, and somehow the ad they made for it raises even more questions than the product itself. The founder has described Friend as a kind of “confidant, friend, God,” says the first 5,000 devices sold out, and now wants to sell 50,000 of these. And there are clearly people who genuinely want this. One commenter on X even said AI makes a better counselor than therapy. So… am I being old-

15h ago
The Decoder

Claude Code now runs daily maintenance on Anthropic's software with a 46 percent merge rate

Anthropic is testing whether Claude Code can handle daily maintenance of the company's own apps, from crash fuzzing to dead-code removal. In a few weeks, the AI created 388 pull requests, and 46 percent were merged after human review. Claude Code inventor Boris Cherny sees this as "early signs of life that this might be possible." The article Claude Code now runs daily maintenance on Anthropic's software with a 46 percent merge rate appeared first on The Decoder .

18h ago
Hugging Face daily papers Free tool

CW-BASS v2: Saturation-Aware Pseudo-Label Selection for Semi-Supervised Segmentation under Foundation-Model Teachers

Semi-supervised semantic segmentation has long turned on one question, which pseudo-labels to trust, and a generation of selection rules, dynamic thresholds, per-class curricula, soft confidence weights, answered it for the noisy, under-confident ResNet teachers of their day. Self-supervised foundation encoders change the regime: with a DINOv2 teacher, confidence saturates, so the filtering that helped a weak teacher can hurt a strong one. We propose CW-BASS v2, a saturation-aware pseudo-label selection method that reads the teacher's confidence regime rather than committing to one rule. It pa

2d ago
Hugging Face daily papers Free tool

DreamX-Phi 1.0: Action-Conditioned Video World Model for Robotic Manipulation

We present DreamX-Phi 1.0, an action-conditioned video world model for robotic manipulation that, given an observed frame, a language instruction, and a prescribed action sequence comprising end-effector poses and gripper states, predicts the resulting future observations. Yet realism alone does not guarantee faithfulness: a convincing rollout can still move the wrong arm or lose the manipulated object. To ensure the prediction respects each arm's commanded path, we inject per-arm SE(3) transformations into attention via PRoPE-style geometric encoding, preserving arm identity and rigid-motion

2d ago
Hugging Face daily papers Free tool

Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation

Interactive autoregressive video generation demands both low-latency rollouts and precise online control. Few-step distillation accelerates generation by reducing denoising steps, while online control imposes a causal constraint: frames and blocks should depend on history and controls available during generation. Existing video distribution matching distillation (DMD) pipelines, however, often supervise causal few-step students using bidirectional teachers that score complete clips. The score for a target can therefore depend on future frames and controls that were unavailable when the student

2d ago
Hugging Face daily papers Free tool

TailBooster: A Dual-Layer Generative Framework for Extreme Value Augmentation with Operational Validity Enforcement

Extreme events in air transport, such as severe arrival delays and abnormal air times, cause cascading network disruptions with substantial operational, economic, and safety costs. Such events are rare in historical records, leaving insufficient training signal for machine learning models. Synthetic data augmentation offers a principled solution, but conventional generative models under-represent distributional tails and give no guarantee against operationally infeasible instances, such as a short air time paired with a long flight distance. No existing approach addresses both limitations for

3d ago
Hugging Face daily papers Free tool

Mitigating Gender Bias in English to Romanian Machine Translation

Machine translation (MT) systems often fail to correctly translate gender, especially when converting from a gender-neutral language like English to a gendered target language such as Romanian. This bias results in translations that default to masculine forms or reinforce gender stereotypes. We propose a hybrid pipeline to mitigate this issue by combining large language model (LLM)-based gender classification with neural machine translation (NMT). Our system uses a fine-tuned LLM to detect the intended gender of target words in English sentences and insert inline gender hint tags. These tagged

6d ago
Hugging Face daily papers Free tool

LLMRouter: Unified Infrastructure for Developing, Evaluating, and Deploying LLM Routers

No single large language model (LLM) is optimal across all queries and budget constraints, making model routing essential for cost-effective deployment. Existing routers adopt diverse formulations and implementations, making fair comparison and extension difficult. We present a unified formulation of LLM routing as a sequential decision process characterized by five components: context encoders, model encoders, scoring functions, decision rules, and learning signals, covering single-turn, multi-turn, and personalized routing. Based on this formulation, we develop an automated pipeline for cons

Aug 7
Hugging Face daily papers Free tool

Maglev: Sliding Recurrent Memory

We introduce , a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training. consists of two coupled models: a prefiller Q, which leverages full attentionIn practice, we use interleaved full and sliding-window attention for Q, as this yields stronger performance. The essential requirement is that Q be more expressive than P, with access to the full history. to produce memory targets m'_t, and a decoder P, which uses only sliding-window attention and recurrent K/V injection to produce decoder memories m_t f

Aug 5
Hugging Face daily papers Free tool

Knowing When to Quit: Diagnosing and Training LLMs to Abort Futile Reasoning

Large language models generate computationally expensive yet semantically void reasoning on beyond-capability tasks, creating risks where plausible-sounding but incorrect derivations mislead users. We characterize this futile reasoning phenomenon through systematic analysis, revealing universal capability overreach and systematic miscalibration between capability and behavior. The dominant failure mode is specious reasoning, which outputs look superficially valid but contain subtle errors, escalating with task difficulty. To address this, we introduce CaRL (Capability-aligned Reinforcement Lea

Jul 31

Watchlist

Posts from the accounts I follow on X.

Scholars

Nothing from this group today.

Daily Curators

@alliekmiller Allie K. Miller

@bentossell Ben Tossell

@dair_ai DAIR AI

@mreflow Matt Wolfe

@rowancheung Rowan Cheung

@TheRundownAI The Rundown AI

Executives & Thought Leaders

@AndrewYNg Andrew Ng

@c_valenzuelab Cristóbal Valenzuela

@elonmusk Elon Musk

@EMostaque Emad Mostaque

@gdb Greg Brockman

@natolambert Nathan Lambert

Official Accounts

@AdobeFirefly Adobe Firefly

@AnthropicAI Anthropic

@NVIDIAAI NVIDIA AI

@pika_labs Pika

@runwayml Runway

The Kitchen

What ran today and what it cost.

Today's spend $0.0570 8 calls · 36,318 tokens
Last 30 days $1.82 266 calls · 2,502,858 tokens
Full breakdown Open The Kitchen → Per-provider rollups, sparklines, model registry