Skip to Main Content

AI Prompt Playground

Multi-Model LLM Testing Environment

Ready to test

Select models from the top right, write your prompt, and click Run to compare outputs side-by-side.

Sampling 5 chain-of-thought paths and majority-voting the answer improves accuracy by 12-18% over single-path CoT on ari.Wang et al., 'Self-Consistency Improves Chain of T…