AINews Portal

Research Command Center

One ranked queue for papers, model launches, repos, agent installs, and cross-market opportunities. JSON is available at /api/research.

Opportunity

12

Paper

20

Model

20

Repo

20

Agent

19

Repo
100/100★ 118.1K · +406/24h
Shubhamsaboo/awesome-llm-apps is moving on GitHub

Check whether this repo maps to a tracked tool, model, or implementation note.

Open signal
1
Repo
100/100GitHub repos★ 118.1K · +406/24h1mo
Shubhamsaboo/awesome-llm-apps is moving on GitHub

100+ AI Agent & RAG apps you can actually run — clone, customize, ship.

Check whether this repo maps to a tracked tool, model, or implementation note.

Stars: 118,09224h stars: 4067d stars: 1,639Credibility: organic: Fork ratio and recent velocity look consistent with normal open-source attention.
2
Repo
100/100GitHub repos★ 228.7K · +356/24h1mo
affaan-m/ECC is moving on GitHub

The agent harness performance optimization system. Skills, instincts, memory, security, and research-first development for Claude Code, Codex, Opencode, Cursor and beyond.

Check whether this repo maps to a tracked tool, model, or implementation note.

Stars: 228,68124h stars: 3567d stars: 2,570Credibility: organic: Fork ratio and recent velocity look consistent with normal open-source attention.
3
Opportunity
97/100Jobs66 roles5h
Anthropic is hiring aggressively

66 open AI roles in the current jobs snapshot.

Use the hiring pattern to infer priorities, buyer pain, and talent-market momentum.

Open roles: 66Newest post: 2026-08-24Company page: not linked yet
4
Opportunity
97/100Jobs29 roles5h
Databricks is hiring aggressively

29 open AI roles in the current jobs snapshot.

Use the hiring pattern to infer priorities, buyer pain, and talent-market momentum.

Open roles: 29Newest post: 2026-08-24Company page: not linked yet
5
Model
97/100Model releaseschat3d
Ox Alpha is a new model to evaluate

stealth released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: stealthUse case: chatContext: 1MInput price: $0.00/M
6
Opportunity
97/100ReleasesElo 144110d
Qwen3.8 27B launched

qwen released a chat model with Elo 1441.

Evaluate benchmark movement, pricing impact, and whether tool pages should reference it.

Organization: qwenUse case: chatElo: 1,441Released: 2026-08-14
7
Repo
97/100GitHub repos★ 80.9K · +590/24h1mo
DietrichGebert/ponytail is moving on GitHub

Makes your AI agent think like the laziest senior dev in the room. The best code is the code you never wrote.

Review credibility signals before using this repo as proof of adoption.

Stars: 80,87724h stars: 5907d stars: 6,783Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
8
Repo
97/100GitHub repos★ 82.5K · +501/24h1mo
Graphify-Labs/graphify is moving on GitHub

AI coding assistant skill (Claude Code, Codex, OpenCode, Cursor, Gemini CLI, and more). Turn any folder of code, SQL schemas, R scripts, shell scripts, docs, papers, images, or videos into a queryable knowledge graph. App code + database schema + infrastructure in one graph.

Review credibility signals before using this repo as proof of adoption.

Stars: 82,54724h stars: 5017d stars: 4,796Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
9
Repo
97/100GitHub repos★ 213.4K · +494/24h1mo
NousResearch/hermes-agent is moving on GitHub

The agent that grows with you

Review credibility signals before using this repo as proof of adoption.

Stars: 213,38024h stars: 4947d stars: 4,063Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
10
Repo
97/100GitHub repos★ 149.5K · +485/24h1mo
firecrawl/firecrawl is moving on GitHub

The API to search, scrape, and interact with the web at scale. 🔥

Review credibility signals before using this repo as proof of adoption.

Stars: 149,49624h stars: 4857d stars: 4,950Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
11
Repo
97/100GitHub repos★ 30.2K · +436/24h1mo
DeusData/codebase-memory-mcp is moving on GitHub

High-performance code intelligence MCP server. Indexes codebases into a persistent knowledge graph — average repo in milliseconds. 158 languages, sub-ms queries, 99% fewer tokens. Single static binary, zero dependencies.

Review credibility signals before using this repo as proof of adoption.

Stars: 30,23524h stars: 4367d stars: 3,940Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
12
Repo
97/100GitHub repos★ 55K · +435/24h1mo
Panniantong/Agent-Reach is moving on GitHub

Give your AI agent eyes to see the entire internet. Read & search Twitter, Reddit, YouTube, GitHub, Bilibili, XiaoHongShu — one CLI, zero API fees.

Review credibility signals before using this repo as proof of adoption.

Stars: 55,01324h stars: 4357d stars: 4,237Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
13
Repo
97/100GitHub repos★ 77.4K · +431/24h1mo
addyosmani/agent-skills is moving on GitHub

Production-grade engineering skills for AI coding agents.

Review credibility signals before using this repo as proof of adoption.

Stars: 77,39924h stars: 4317d stars: 8,435Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
14
Repo
97/100GitHub repos★ 56.3K · +423/24h1mo
asgeirtj/system_prompts_leaks is moving on GitHub

Extracted system prompts from Anthropic - Claude Fable 5, Opus 4.8, Claude Code, Claude Design. OpenAI - ChatGPT GPT-5.6, Codex GPT-5.6, GPT-5.5. Google - Gemini 3.5 Flash, 3.1 Pro, Antigravity. xAI - Grok, Cursor, Copilot, VS Code, Perplexity, and more. Updated regularly.

Review credibility signals before using this repo as proof of adoption.

Stars: 56,34024h stars: 4237d stars: 7,241Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
15
Repo
97/100GitHub repos★ 88.2K · +395/24h1mo
JuliusBrussee/caveman is moving on GitHub

🪨 why use many token when few token do trick — Claude Code skill that cuts 65% of tokens by talking like caveman

Review credibility signals before using this repo as proof of adoption.

Stars: 88,23224h stars: 3957d stars: 4,052Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
16
Repo
97/100GitHub repos★ 37.2K · +380/24h1mo
calesthio/OpenMontage is moving on GitHub

World's first open-source, agentic video production system. 12 pipelines, 52 tools, 500+ agent skills. Turn your AI coding assistant into a full video production studio.

Review credibility signals before using this repo as proof of adoption.

Stars: 37,20224h stars: 3807d stars: 3,972Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
17
Repo
97/100GitHub repos★ 7.8K · +371/24h1mo
wonderwhy-er/DesktopCommanderMCP is moving on GitHub

This is MCP server for Claude that gives it terminal control, file system search and diff file editing capabilities

Review credibility signals before using this repo as proof of adoption.

Stars: 7,82724h stars: 3717d stars: 1,582Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
18
Repo
97/100GitHub repos★ 130.5K · +362/24h1mo
msitarzewski/agency-agents is moving on GitHub

A complete AI agency at your fingertips - From frontend wizards to Reddit community ninjas, from whimsy injectors to reality checkers. Each agent is a specialized expert with personality, processes, and proven deliverables.

Review credibility signals before using this repo as proof of adoption.

Stars: 130,54324h stars: 3627d stars: 3,416Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
19
Repo
97/100GitHub repos★ 7.1K · +321/24h1mo
google-labs-code/stitch-skills is moving on GitHub

A library of Agent Skills designed to work with the Stitch MCP server. Each skill follows the Agent Skills open standard, for compatibility with coding agents such as Antigravity, Gemini CLI, Claude Code, Cursor.

Review credibility signals before using this repo as proof of adoption.

Stars: 7,13624h stars: 3217d stars: 762Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
20
Repo
97/100GitHub repos★ 191K · +305/24h1mo
multica-ai/andrej-karpathy-skills is moving on GitHub

A single CLAUDE.md file to improve Claude Code behavior, derived from Andrej Karpathy's observations on LLM coding pitfalls.

Review credibility signals before using this repo as proof of adoption.

Stars: 190,98124h stars: 3057d stars: 3,293Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
21
Repo
97/100GitHub repos★ 27.8K · +246/24h1mo
JCodesMore/ai-website-cloner-template is moving on GitHub

Clone any website with one command using AI coding agents

Review credibility signals before using this repo as proof of adoption.

Stars: 27,82424h stars: 2467d stars: 2,199Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
22
Repo
97/100GitHub repos★ 70.4K · +225/24h1mo
rtk-ai/rtk is moving on GitHub

CLI proxy that reduces LLM token consumption by 60-90% on common dev commands. Single Rust binary, zero dependencies

Review credibility signals before using this repo as proof of adoption.

Stars: 70,40924h stars: 2257d stars: 1,855Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
23
Repo
97/100GitHub repos★ 77.4K · +211/24h1mo
nexu-io/open-design is moving on GitHub

🎨 The open-source Claude Design alternative. 🖥️ Local-first desktop app. 🖼️ Your coding agent becomes the design engine: prototypes, landing pages, dashboards, slides, images & video — real files, HTML/PDF/PPTX/MP4 export. 🤖 Claude Code / Codex / Cursor / Gemini / OpenCode / Qwen & 20+ CLIs via BYOK.

Review credibility signals before using this repo as proof of adoption.

Stars: 77,39924h stars: 2117d stars: 2,384Credibility: velocity: Recent star velocity is unusually high; inspect source quality before trusting rank.
24
Opportunity
96/100Jobs16 roles18h
xAI is hiring aggressively

16 open AI roles in the current jobs snapshot.

Use the hiring pattern to infer priorities, buyer pain, and talent-market momentum.

Open roles: 16Newest post: 2026-08-23Company page: not linked yet
25
Opportunity
84/100ReleasesElo 137613d
Solar Pro 4 launched

upstage released a chat model with Elo 1376.

Evaluate benchmark movement, pricing impact, and whether tool pages should reference it.

Organization: upstageUse case: chatElo: 1,376Released: 2026-08-10
26
Model
83/100Model releasesElo 144110d
Qwen3.8 27B is a new model to evaluate

qwen released a chat model with Elo 1441.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: qwenUse case: chatContext: 262.1KInput price: $0.40/M
27
Model
83/100Model releaseschat10d
GLM 5.3 is a new model to evaluate

z-ai released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: z-aiUse case: chatContext: 1MInput price: $1.40/M
28
Model
82/100Model releaseschat11d
Gemini 3.7 Flash is a new model to evaluate

google released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: googleUse case: chatContext: 1MInput price: $0.38/M
29
Model
73/100Model releaseschat2d
Muse Spark 1.2 Contributor is a new model to evaluate

meta released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: metaUse case: chatContext: 1MInput price: $0.10/M
30
Model
73/100Model releasesvision3d
DeepSeek V4 Flash Vision Exp is a new model to evaluate

deepseek released a vision model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: deepseekUse case: visionContext: 1MInput price: $0.22/M
31
Opportunity
72/100Releaseschat2d
Muse Spark 1.2 Contributor launched

meta released a chat model.

Evaluate benchmark movement, pricing impact, and whether tool pages should reference it.

Organization: metaUse case: chatElo: not availableReleased: 2026-08-21
32
Model
72/100Model releaseschat4d
Hy-MT2-1.8B is a new model to evaluate

tencent released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: tencentUse case: chatContext: 8.2KInput price: $0.04/M
33
Model
72/100Model releaseschat4d
Hy-MT2-30B-A3B is a new model to evaluate

tencent released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: tencentUse case: chatContext: 8.2KInput price: $0.07/M
34
Model
72/100Model releaseschat11d
DeepSeek V4 Pro 0813 is a new model to evaluate

deepseek released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: deepseekUse case: chatContext: 1MInput price: $1.12/M
35
Model
72/100Model releaseschat12d
Grok 4.6 is a new model to evaluate

x-ai released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: x-aiUse case: chatContext: 500KInput price: $2.00/M
36
Opportunity
71/100Pricing-46.7%8h
Qwen: Qwen3.6 27B got cheaper

qwen input per m price changed from $0.60/M to $0.32/M.

Re-check workloads in the optimizer and consider switching high-volume traffic.

Provider: qwenPrice field: input per mOld price: $0.60/MNew price: $0.32/M
37
Opportunity
71/100Releasesvision3d
DeepSeek V4 Flash Vision Exp launched

deepseek released a vision model.

Evaluate benchmark movement, pricing impact, and whether tool pages should reference it.

Organization: deepseekUse case: visionElo: not availableReleased: 2026-08-21
38
Model
71/100Model releaseschat4d
GLM Latest is a new model to evaluate

z-ai released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: z-aiUse case: chatContext: 1MInput price: $1.40/M
39
Model
71/100Model releaseschat4d
Hy-MT2-7B is a new model to evaluate

tencent released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: tencentUse case: chatContext: 8.2KInput price: $0.07/M
40
Model
71/100Model releaseschat13d
Nemotron 3.5 Lightning is a new model to evaluate

nvidia released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: nvidiaUse case: chatContext: 262.1KInput price: $0.08/M
41
Model
71/100Model releasesElo 137613d
Solar Pro 4 is a new model to evaluate

upstage released a chat model with Elo 1376.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: upstageUse case: chatContext: 524.3KInput price: $0.03/M
42
Model
70/100Model releasesimage14d
Muse Glimmer 30B is a new model to evaluate

meta released a image model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: metaUse case: imageContext: 131.1KInput price: $0.35/M
43
Opportunity
69/100Releaseschat3d
Ox Alpha launched

stealth released a chat model.

Evaluate benchmark movement, pricing impact, and whether tool pages should reference it.

Organization: stealthUse case: chatElo: not availableReleased: 2026-08-20
44
Opportunity
69/100Releaseschat4d
Hy-MT2-1.8B launched

tencent released a chat model.

Evaluate benchmark movement, pricing impact, and whether tool pages should reference it.

Organization: tencentUse case: chatElo: not availableReleased: 2026-08-20
45
Opportunity
69/100Releaseschat4d
Hy-MT2-30B-A3B launched

tencent released a chat model.

Evaluate benchmark movement, pricing impact, and whether tool pages should reference it.

Organization: tencentUse case: chatElo: not availableReleased: 2026-08-20
46
Model
68/100Model releaseschat10d
Dots3-Note Preview (free) is a new model to evaluate

dots-studio released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: dots-studioUse case: chatContext: 512KInput price: unknown
47
Opportunity
67/100Pricing-42.9%8h
DeepSeek V4 Flash Latest got cheaper

deepseek output per m price changed from $0.14/M to $0.08/M.

Re-check workloads in the optimizer and consider switching high-volume traffic.

Provider: deepseekPrice field: output per mOld price: $0.14/MNew price: $0.08/M
48
Model
67/100Model releaseschat11d
Seed 2.1 Turbo is a new model to evaluate

bytedance-seed released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: bytedance-seedUse case: chatContext: 262.1KInput price: $0.50/M
49
Model
67/100Model releaseschat11d
Qwen3.8 2.4T A95B is a new model to evaluate

qwen released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: qwenUse case: chatContext: 1MInput price: $2.00/M
50
Repo
67/100GitHub repos★ 26.1K · +196/24h1mo
alibaba/page-agent is moving on GitHub

JavaScript in-page GUI agent. Control web interfaces with natural language.

Review credibility signals before using this repo as proof of adoption.

Stars: 26,08624h stars: 1967d stars: 0Credibility: stale: Large historical star count with little recent movement.
51
Model
66/100Model releaseschat12d
LFM2.5-2.6B (free) is a new model to evaluate

liquid released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: liquidUse case: chatContext: 128KInput price: unknown
52
Model
66/100Model releaseschat13d
Sakana Namazu is a new model to evaluate

sakana released a chat model.

Compare benchmark fit, context, pricing, and whether it belongs in model recommendations.

Organization: sakanaUse case: chatContext: 262.1KInput price: $0.95/M
53
Paper
55/100Research papers0 mentions · ▲ 2574d
EnvHarness: Awakening Static Worlds for Agent Learning

EnvHarness and EnvRigger dynamically reshape static environments via programmable plugins to target agent weaknesses and improve reinforcement learning co-evolution.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.1988048h mentions: 07d mentions: 0HF upvotes: 257Implementation: not detected
54
Paper
51/100Research papers0 mentions · ▲ 3749d
StateM: Reaching 95.3% Raw Accuracy, or a \$15 Frontier Run, on Terminal-Bench 2.1 via Harness Scaling

StateM is a runtime system that improves long-horizon agent execution through durable states, recoverable runbooks, and enforceable procedural controls without altering model weights.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.1508948h mentions: 07d mentions: 0HF upvotes: 374Implementation: not detected
55
Paper
50/100Research papers0 mentions · ▲ 25910d
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination

A new benchmark for AI-generated video detection reveals that current detectors fail to generalize across realistic crisis-related videos and become less reliable as content spreads socially.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.1439148h mentions: 07d mentions: 0HF upvotes: 259Implementation: not detected
56
Paper
48/100Research papers0 mentions · ▲ 24512d
Spark-to-Paper: End-to-End Research Paper Generation as a Composable Skill

Spark-to-Paper is a lightweight, composable workflow inside coding assistants that generates research papers by separating planning from reporting, enforcing evidence-based claim revision, and using integrity checks to reduce fabrication.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.1192448h mentions: 07d mentions: 0HF upvotes: 245Implementation: not detected
57
Paper
48/100Research papers0 mentions · ▲ 18113d
ComBodied Agents: a New Paradigm of Human-Centric Agentic AI

Combodied Agents integrate digital and embodied tools into a closed-loop framework that models individual human-state trajectories over time to provide proportionate, consent-aware support.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.1091548h mentions: 07d mentions: 0HF upvotes: 181Implementation: not detected
58
Paper
48/100Research papers0 mentions · ▲ 17113d
Beyond Pixels: From Video Priors to 4D Worlds

Latent-to-4D enables reusable direct 4D generation from video diffusion latents via alignment with a pretrained decoder and spatiotemporal attention, transferring across generators without retraining.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.1074448h mentions: 07d mentions: 0HF upvotes: 171Implementation: not detected
59
Paper
47/100Research papers0 mentions · ▲ 55314d
BDH-CQ: In-Context Learning with Recurrent Latent Reasoning

A 150M-parameter reasoning model using recurrent latent reasoning and in-context learning achieves a new cost-accuracy frontier on ARC-AGI-1.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.0988848h mentions: 07d mentions: 0HF upvotes: 553Implementation: not detected
60
Paper
47/100Research papers0 mentions · ▲ 30314d
Macaron-V1: Towards Open Continual Learning with Self-Improvement and Mixture-of-LoRA

Macaron-V1 is an open agent-model family that uses a Mixture-of-LoRA architecture and recursive self-improvement to enable continual learning and collaboration across specialized tasks.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.0981948h mentions: 07d mentions: 0HF upvotes: 303Implementation: not detected
61
Paper
47/100Research papers0 mentions · ▲ 18715d
On-Policy Self-Distillation without Any Supervision

Unsupervised on-policy self-distillation improves large language models by using internal consistency and majority-vote pseudo-solutions to correct confident errors without external supervision.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.0629648h mentions: 07d mentions: 0HF upvotes: 187Implementation: not detected
62
Paper
44/100Research papers0 mentions · ▲ 22319d
Recursive Synthesis for Long-Horizon Terminal Tasks

This paper introduces a method to automatically generate large volumes of training data for AI agents that perform long, multi-step tasks in terminal environments (like running commands in a shell). The key problem is that such data is normally hand-crafted and costs hundreds to thousands of dollars per task because the instruction, environment, and solution must all be consistent. The new approach, called Recursive Synthetic Terminal Tasks (RST), starts with a few verified tasks and repeatedly extends their solutions, then automatically updates the instructions and checks to match, validating each new task in a sandbox. Over 15 rounds, it produced 37,484 tasks at about $0.05 each, with difficulty rising sharply—median solution length grew from 67 to 374 lines, and the best model's success rate dropped from 90% to 2.5%. Training models on these synthetic tasks improved their performance on standard terminal-agent benchmarks by up to 41% relative, and the process shows no signs of plateauing, suggesting it could scale further.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.0546648h mentions: 07d mentions: 0HF upvotes: 223Implementation: not detected
63
Paper
43/100Research papers0 mentions · ▲ 25023d
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution

OpenART introduces a scalable red-teaming arena with evolving stateful environments to evaluate long-horizon AI agent safety, using the EMHA attack policy to expose increasing failure rates as task complexity grows.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2608.0067748h mentions: 07d mentions: 0HF upvotes: 250Implementation: not detected
64
Paper
42/100Research papers0 mentions · ▲ 29425d
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents

This paper introduces Qwen-UI-Agent, a foundation model designed to act as a general-purpose executor across mobile, desktop, web, and search interfaces. Unlike typical GUI agents that operate in isolated sandboxes, it runs on real devices and unifies GUI clicks with command-line actions, generating multiple steps per turn to handle long, complex workflows. A key innovation is an automated data flywheel where the agent itself builds tasks, diagnoses its own failures, and plans improvements, plus online reinforcement learning on trajectories over 100 steps using 10,000+ parallel environments. The headline results show state-of-the-art mobile performance (82.1% on MobileWorld, 92.2% on MobileWorld-Real) and competitive computer/browser scores (79.5% on OSWorld-Verified, 73.6% on WebArena) against frontier models like GPT-5.6 and Gemini 3.1 Pro. The work matters because it pushes agents from scripted demos toward reliable, self-improving operation on actual devices, a prerequisite for practical digital assistants.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2607.2822748h mentions: 07d mentions: 0HF upvotes: 294Implementation: not detected
65
Paper
42/100Research papers0 mentions · ▲ 29325d
AskChem: Claim-Centered Infrastructure for Chemistry Literature Synthesis

AskChem changes chemistry literature search from returning ranked papers to returning atomic, provenance-carrying claims—each tied to a source DOI and a verbatim quote or evidence locator. It indexes 2.4M claims from 147K papers and offers multiple access paths: a faceted taxonomy for browsing, an evidence graph linking claims, and a living taxonomy organized by scientific principles, plus REST, SDK, and MCP interfaces for AI agents. In their benchmark, grounding a GPT-5.5 reader in AskChem achieved 100% resolvable DOIs versus 88.3% without retrieval, and the highest citation density among five systems tested. This matters because it directly addresses the manual assembly and verification burden in cross-paper synthesis, for both humans and AI agents.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2607.2861848h mentions: 07d mentions: 0HF upvotes: 293Implementation: not detected
66
Paper
42/100Research papers0 mentions · ▲ 16925d
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering

This paper introduces OpenMLE, an open-source framework for studying recursive self-improvement (RSI) in machine learning engineering, where AI systems help build better AI. The framework includes a verifiable task environment, reinforcement learning for operator learning, and long-horizon evolutionary search. Using this stack, the authors post-train a 35B-parameter model called Frontis-MA1 to act as a meta-evolution agent, composing four atomic program-editing operations—Draft, Improve, Debug, and Crossover—into an iterative loop. On the MLE-Bench Lite benchmark, the model improves its Medal Average from 39.39% to 60.61% over its base model, and reaches 71.21% with an enhanced search variant, outperforming GPT-5.5 + Codex and approaching much larger models like GPT-5.6 Sol and Kimi K3. The results demonstrate that both the trained model and the search framework independently contribute to performance gains, and the authors release the model weights and code to enable reproducible RSI research.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2607.2856848h mentions: 07d mentions: 0HF upvotes: 169Implementation: not detected
67
Paper
41/100Research papers0 mentions · ▲ 33628d
Kimi K3: Open Frontier Intelligence

Kimi K3 is a 2.8-trillion-parameter open-source AI model that uses a Mixture-of-Experts design to activate only 104 billion parameters per token, making it efficient despite its size. It introduces new attention mechanisms and a stable expert-routing method to improve performance across long sequences and deep networks, achieving roughly 2.5x better scaling efficiency than its predecessor. The model excels at long-horizon coding, agentic tasks, reasoning, and vision, outperforming all other open models and most proprietary ones, though it still trails the top closed models like Claude Fable 5 and GPT-5.6 Sol. Its release as open weights aims to accelerate research and deployment of frontier-level AI.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2607.2465348h mentions: 07d mentions: 0HF upvotes: 336Implementation: not detected
68
Paper
40/100Research papers0 mentions · ▲ 2031mo
ABot-World-0: Infinite Interactive World Rollout on a Single Desktop GPU

ABot-World-0 is a video world model that runs real-time, interactive simulations on a single desktop GPU. It learns to generate coherent video sequences in response to keyboard actions by training on a diverse mix of AAA games, simulations, and internet videos. The key innovation is a training method called LongForcing that aligns the model's long self-generated sequences with a more accurate teacher model, preventing the drift that typically plagues long video generation. On a single RTX 5090, it streams 720P video at up to 16 FPS with only 1.2 seconds of latency, enabling interactive scene roaming and character control.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2607.1919148h mentions: 07d mentions: 0HF upvotes: 203Implementation: not detected
69
Paper
39/100Research papers0 mentions · ▲ 1831mo
RynnBrain 1.1: Towards More Capable and Generalizable Embodied Foundation Model

RynnBrain 1.1 is a family of embodied AI models (2B, 9B, and a 122B-A10B mixture-of-experts variant) trained to handle perception, spatial reasoning, localization, and planning for robots. The key upgrades over version 1.0 are contact-point prediction (where a robot should touch an object) and native 3D grounding for the smaller models, which makes their outputs more directly usable for manipulation. The authors also built a separate VLA (vision-language-action) model with a shared action space across different robot bodies, using masking to handle their differing morphologies, and tested it on three real robots. On benchmarks like VSI-Bench and RefSpatial-Bench, the largest model beats all proprietary and open-source competitors, and real-robot tests show that policies initialized with RynnBrain outperform those based on Qwen or other generalist VLAs. Notably, training on multiple tasks and robot types simultaneously improved both process scores and success rates compared to training each task separately.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2607.1797748h mentions: 07d mentions: 0HF upvotes: 183Implementation: not detected
70
Paper
39/100Research papers0 mentions · ▲ 1821mo
LongStraw: Long-Context RL Beyond 2M Tokens under a Fixed GPU Budget

LongStraw is a system that enables reinforcement learning (RL) post-training on contexts exceeding 2 million tokens using a fixed GPU budget, addressing the gap where inference handles million-token contexts but RL training typically tops out at 256K tokens. It works by evaluating the shared prompt without autograd, keeping only essential model state for later tokens, and replaying short response branches one at a time—trading extra replay time for a smaller live training graph. Implemented on Qwen3.6-27B and GLM-5.2 models, LongStraw completed grouped scoring and backward passes at 2.1 million positions on eight H20 GPUs, with a stress test reaching 4.46 million positions. This matters for AI agents that accumulate long trajectories of observations and tool outputs, though the paper notes these experiments establish execution capacity rather than full training correctness.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2607.1495248h mentions: 07d mentions: 0HF upvotes: 182Implementation: not detected
71
Paper
38/100Research papers0 mentions · ▲ 1871mo
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable

Modern AI agents rely on a 'harness'—the code that builds prompts, manages state, calls tools, and coordinates execution—which must be constantly updated as models and APIs change. The main bottleneck is finding exactly which lines of code implement a desired behavior, because harness code is large, tangled, and spread across many files. This paper introduces the Harness Handbook, an automatically generated document that maps each high-level behavior to its source code locations, and a method called Behavior-Guided Progressive Disclosure (BGPD) that helps both human developers and coding agents navigate from a behavior description down to the relevant code. In tests on two real open-source harnesses, the handbook approach improved the accuracy of locating code and planning edits while using fewer tokens, especially for behaviors scattered across modules or on rarely executed paths.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2607.1328548h mentions: 07d mentions: 0HF upvotes: 187Implementation: not detected
72
Paper
37/100Research papers0 mentions · ▲ 1871mo
Orca: The World is in Your Mind

Orca establishes a unified world latent space through next-state-prediction modeling using multimodal data and demonstrates superior performance in downstream tasks compared to specialized baselines.

Read the abstract and check whether the method changes tool, benchmark, or model coverage.

arXiv: 2606.3053448h mentions: 07d mentions: 0HF upvotes: 187Implementation: not detected
73
Agent
20/100Agent skillsrepo config · 0 heat
template Skill is worth testing

Agent skill from the official anthropics/skills repository.

Try the install path in a sandbox and decide whether it should enter the operator toolkit.

Category: Agent skillMentions: 0Install: repo-configRisk: medium
74
Agent
10/100MCP registryremote ready · 0 heat
ComplyHat MCP is worth testing

MCP for AI compliance documentation across SR 26-2, EU AI Act, NIST AI RMF, and ISO/IEC 42001.

Try the install path in a sandbox and decide whether it should enter the operator toolkit.

Category: MCP serverMentions: 0Install: remote-readyRisk: lower
75
Agent
10/100MCP registryremote ready · 0 heat
Demand Discovery AI MCP is worth testing

Sales-agent MCP for Demand Discovery AI. Validate startup ideas with behavioral signals.

Try the install path in a sandbox and decide whether it should enter the operator toolkit.

Category: MCP serverMentions: 0Install: remote-readyRisk: lower
76
Agent
10/100MCP registryremote ready · 0 heat
DiagramZu MCP is worth testing

Let Claude, Cursor, or ChatGPT author Mermaid diagrams your team can read and share.

Try the install path in a sandbox and decide whether it should enter the operator toolkit.

Category: MCP serverMentions: 0Install: remote-readyRisk: lower
77
Agent
10/100MCP registryremote ready · 0 heat
DoneThat MCP is worth testing

Privacy-first work tracking with summaries, reports, coaching, and AI-ready long-term memory.

Try the install path in a sandbox and decide whether it should enter the operator toolkit.

Category: MCP serverMentions: 0Install: remote-readyRisk: lower
78
Agent
10/100MCP registryremote ready · 0 heat
Dreamlit MCP is worth testing

Create, test, publish, and manage Dreamlit notification workflows from AI clients.

Try the install path in a sandbox and decide whether it should enter the operator toolkit.

Category: MCP serverMentions: 0Install: remote-readyRisk: lower
79
Agent
10/100MCP registryremote ready · 0 heat
Filtrix AI MCP MCP is worth testing

Filtrix MCP for image/video generation. Portal: https://agent.filtrix.ai/

Try the install path in a sandbox and decide whether it should enter the operator toolkit.

Category: MCP serverMentions: 0Install: remote-readyRisk: lower
80
Agent
10/100MCP registryremote ready · 0 heat
fincontext MCP is worth testing

Read-only bank & investment accounts via Plaid: balances, holdings, transactions, SQL analytics.

Try the install path in a sandbox and decide whether it should enter the operator toolkit.

Category: MCP serverMentions: 0Install: remote-readyRisk: lower

Scores are synthesized from the portal's own tables: Radar, papers, model releases, repo snapshots, and MCP/skill discovery. Use this as an editorial triage queue, then verify primary sources before publishing.