Have been looking at the documentation for this, I think I have myself a weekend side project. The framework has interesting tools.
Amongst them is evaluation as well, I think its trying to bring engineering disipline to skill creation.
What I am thinking is can we create an iterative workflow for building agent skills: draft a skill, evaluate it with SkillEvaluator (especially the live Skill Lift measurements), then use the results to improve it. Loop until the skill is high-quality, safe, and demonstrably helpful.
Maybe create a workflow for Atomic.
Wonder if we could use DSPy to optimize the prompt, could we use SkillEvaluator to generate enough synthetic, high quality eval data...
Mind just goes. Just wanted to share