115 tools found
A read-only skill that scores installed skills across 8 research-grounded dimensions using session data plus static checks.
Five agent skills that evaluate interfaces against 168 research-backed UX/UI principles and detect antipatterns.
Promptfoo example that benchmarks OpenAI gpt-5.4-mini model with two reasoning effort settings (none vs. medium) to compare output quality, latency, and cost.
picklesdev is an autonomous agent for end-to-end workflow automation, enabling users to run complete processes without manual intervention.
Generates the 5 most promising B2C market segments for your product, ranked by attractiveness, using the Advanced JTBD methodology.
Forces a team through a structured positioning exercise to produce an honest, sub-30-word competitive positioning statement.
Bare-bones promptfoo starting example for comparing Anthropic Claude and OpenAI GPT model responses side by side.
Promptfoo example for the conversation-relevance assertion: sliding-window scoring that catches chatbot responses that drift off-topic mid-conversation.
Promptfoo RAG example with intentionally mixed pass/fail context-recall assertions, showing both successful and failing RAG metrics.
Promptfoo example using the search-rubric assertion to verify time-sensitive claims (prices, news, weather) via real-time web search.
Benchmark workflow comparing OpenAI and Azure OpenAI using GPT-5-mini to measure speed, cost, and output differences between the two services.
Promptfoo examples for Perplexity's search-augmented models - citations, structured output, date/location filters, and reasoning modes.