A skill's quality is measurable: does it convert clean, fit the budget, carry a citation, get used, and produce positive outcomes? The platform holds all of that data; this skill turns it into a verdict and an action, not just a report.
Steps
- Call
evaluate_skill {slug}on the uplift MCP server (workspace key required — see uplift-connect). It returns a score, a pass/fail checklist, the conversion report, feedback tallies with recent negative messages, and 30-day workspace usage. - Turn each failed check into its action:
- trigger-style summary → rewrite the description as "Use when …" so agents load it
- over budget → split the skill or cut the weakest section
- conversion residue → tracking or vendor-specific text survives in the served body; clean the source and re-save
- cited → add
source: {repo, path, commit, url}on the next save - fresh → re-read against current practice, then save (a save is the freshness signal, so don't save without actually reviewing)
- used this month → check the trigger description against what the team actually asks for; unused + stale together makes an archive candidate
- positive outcomes → read
recent_negativemessages and fix the instruction they blame
- Route the fix by ownership: workspace skill →
save_skill; public-library skill → change request quoting the evidence (POST /v1/change-requests). - Close the loop: after acting on a skill, report your own outcome with
send_feedback.
Judgement calls
- Zero usage on a skill younger than ~30 days is noise, not failure — flag it, keep it.
- Real outcomes outrank static checks: negative feedback on a checklist-perfect skill still means it fails in practice.
- Archive candidates need all three: unused ≥90 days, stale, and no positive feedback.
Rules
- Quote the numbers and the feedback messages in every recommendation — never "seems fine".
- Evaluate before any public promotion and after any negative-feedback report.
- Batch an audit:
list_skillsfirst, evaluate each, then present one ranked table (healthy / fix / archive) instead of a stream of per-skill verdicts. - Without MCP support, degrade gracefully: fetch the prompt over REST, run its body
through
POST /v1/convert, and judge the static checks locally — say that usage and feedback data were unavailable.