Skip to content
#

plugin-testing

Here are 8 public repositories matching this topic...

Language: All
Filter by language
oolong-pairs

Benchmark harness for A/B testing Claude Code plugins against OOLONG long-context reasoning tasks. Compare truncation vs RLM-RS recursive chunking strategies. Features Claude Code hooks integration, SQLite persistence, and comprehensive scoring aligned with the OOLONG paper methodology.

  • Updated Apr 13, 2026
  • Python

Binary-criteria evaluation harness for Claude skills with planned extension to plugins, agents, and MCP servers. Score every change yes/no across 7 layers — package integrity, trigger quality, functional quality, regression protection, baseline value, model variance, rollout safety. Never gradients.

  • Updated May 7, 2026
  • TypeScript

Improve this page

Add a description, image, and links to the plugin-testing topic page so that developers can more easily learn about it.

Curate this topic

Add this topic to your repo

To associate your repository with the plugin-testing topic, visit your repo's landing page and select "manage topics."

Learn more