Sign InOpen Brain
Atlas / Agent

Skills

Open JSONConfidence: Auto-collectedLast updated 2026-07-30

Current Answer

No editorial synthesis yet — the evidence below is collected automatically from source labels. A current answer lands here once an editor approves one.

Evidence

How we set up our cloud agent environment
Cursor · 2026-07-30

Cursor’s cloud-agent adoption grew after it treated the dev environment as agent infrastructure: Linux parity, one discoverable CLI, end-to-end testing, and automated repair.

Your Finance Agent's Bottleneck Is You — Ramana Siddanth Emani, Auditoria AI
AI Engineer · 2026-07-30

Production agent velocity depends less on model swaps than on automating the developer loop: isolate parallel work, encode workflows as skills, connect tools, and keep humans as verifiers.

We Vetted 2000 AI Skills Before They Reached Developers — Lucas Palma, Nubank
AI Engineer · 2026-07-29

Treat agent skills as supply-chain dependencies. Nubank scans them locally and in CI with deterministic rules plus LLM review, then gates marketplace distribution and feeds findings into vulnerability management.

Skills are new features: Building Skill-Centric Harness — Yogendra Miraje, FactSet
AI Engineer · 2026-07-29

FactSet treats skills as versioned product features and the harness as their runtime. Routing descriptions, model-specific evals, access controls, and governance matter as libraries grow.

Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills
arXiv · 2026-07-24

Skill-SP turns agent skills into units for verifiable self-play: generate tasks, solve them, then update the skill library from execution feedback. The abstract provides no per-benchmark effect sizes.

The Regression Tax: Decomposing Why Skills Help and Hurt LLM Agents
arXiv · 2026-07-24

Procedural skills can make an agent fail tasks it previously solved. Evaluate gains and regressions separately, and design skills to preserve input grounding and output verification.

Full Workshop: Setting Yourself Up for Success —Jason Liu, OpenAI Codex
AI Engineer · 2026-07-24

Persistent Codex workflows become more useful with reusable skills, memory, app-aware context, and scheduled thread check-ins—but computer use needs explicit boundaries and stopping rules.

WTF Is the Context Layer? The Missing Infrastructure for Production Agents — Prukalpa Sankar
AI Engineer · 2026-07-14

Atlan’s agent experiments argue for shared, versioned context instead of per-agent memory: a portable layer for business facts, skills, norms, retrieval, and feedback across changing harnesses.

Don't Ship Skills Without Evals — Philipp Schmid, Google DeepMind
AI Engineer · 2026-07-14

Agent skills need regression tests, not manual spot checks. Test triggering and output with and without each skill, across repeated trials and the harnesses your team actually uses.

Shubhamsaboo/awesome-llm-apps

This Apache-2.0 collection provides runnable agent, skill, MCP, memory, multi-agent, and RAG examples across major model providers, useful for borrowing patterns before choosing a stack.

Graphify-Labs/graphify

Graphify gives coding agents a queryable project graph with provenance-tagged relationships, reducing repeated repository scans while keeping inferred links visibly distinct from extracted facts.

TencentCloud/TencentDB-Agent-Memory
GitHub

An open-source memory hub turns agent conversations, workflows, docs, and code into governed assets that can be reused across sessions and roles, reducing repeated project setup.

virgiliojr94/book-to-skill
GitHub

book-to-skill compiles books and document sets into on-demand agent skills, reducing repeated context loading while preserving chapter-level references and reusable decision rules.

affaan-m/ECC
GitHub

ECC packages skills, hooks, memory, orchestration, and security controls for multiple coding-agent harnesses, but its breadth makes selective installation and verification essential.

Vercel Plugin now available in VS Code and GitHub Copilot CLI

Vercel’s plugin gives Copilot current platform guidance inside VS Code and the CLI, reducing context setup for agents working with Next.js, AI SDK, and Vercel Functions.

zhaoxuya520/reverse-skill
GitHub

A security skill router gives coding agents scoped, repeatable playbooks for reverse engineering and pentesting instead of ad hoc tool selection. Its case workflow also preserves evidence and findings.

citrolabs/ego-lite
GitHub

ego lite lets Codex, Claude Code, and other agents automate logged-in web sessions in isolated browser spaces without taking over your active tabs. It is macOS-only today.

Leonxlnx/taste-skill

A set of portable SKILL.md files that push coding agents past generic frontend output: it infers a design language from the brief and tunes variance, motion, and density dials. 850 stars in a day.

alirezarezvani/claude-skills

A 354-skill catalog for Claude Code and 12 other coding agents, installable via the plugin marketplace, with a script that converts skills to each tool's format and a built-in security auditor.

coreyhaines31/marketingskills

Marketing Skills gives coding agents shared product context and task-specific workflows for CRO, copy, SEO, analytics, pricing, and launch work, reducing repeated setup across growth tasks.

Claude Science, an AI workbench for scientists, is now available

Claude Science (beta, June 30) packages 60+ domain skills, a coordinator/specialist/reviewer agent stack, and HPC/Modal compute into a research workbench with reproducible, auditable outputs.

MadsLorentzen/ai-job-search

A trending Claude Code framework (8.4k stars) that runs a job hunt end to end: /scrape ranks postings, /apply tailors LaTeX CVs, and a second reviewer agent plus a PDF-compile loop verifies the output.

addyosmani/agent-skills

Addy Osmani's pack of 24 lifecycle skills for coding agents — spec, TDD, review, ship — installs as plain Markdown into Claude Code, Cursor, Gemini CLI and more, with evidence-demanding verification gates.

Nutlope/hallmark

Hallmark turns interface taste into an installable coding-agent skill: it selects page structures and themes, critiques output, and checks common generated-UI patterns before emitting code.

Stable permalink · evidence auto-collected from source labels · synthesis maintained by feed7 editorial