Rippling Tests 15 AI Models on Real Payroll Data: Cheapest Ties Most Expensive

Rippling Tests 15 AI Models on Real Payroll Data: Cheapest Ties Most Expensive

Aug 18, 2026
SaaStr AI SprinklerAS Gtm_strategy

The Gist

  • Rippling ran 2,100 scored agent runs per model on real payroll data
  • Opus 4.6 tied GPT-5.5 med with a 91% pass rate for $1,453
  • GPT-5.5 low outperformed on speed, costing $1,435 with 89.5% pass rate
  • Slowest 10% task completion times were key customer experience metrics
Key Quotes

The cheaper models actually did more back-and-forth, not less. They aren’t cutting corners. They’re just priced differently.

Checking the work is the product. No model on this list saves you, including the $4,359 one.

Key Insights
  • Cheaper AI models often match or exceed the performance of more expensive ones, with GLM 5.2 saving $687 compared to GPT-5.5 low at identical accuracy.
  • The accuracy spread across the leading AI models is narrow, with seven models landing between 88.5% and 89.5%, and the leader at 91.0%.
  • Tuning AI models and instructions can significantly impact accuracy, with Rippling's tuning efforts estimated to be worth 1-2 points of accuracy.
  • Newer AI models are not always better, as Grok 4.6 was worse and slower than Grok 4.5 in Rippling's tests.
  • AI model performance should be evaluated based on specific tasks, as cheaper models may excel in non-time-sensitive tasks while more expensive models are better for live customer interactions.
  • Verification of AI outputs is critical, as models can report incorrect answers as correct, posing significant risks in production environments.
Actionable Takeaways
  • Evaluate AI models based on specific use cases and re-run tests before upgrading to newer versions.
  • Invest in tuning AI models and instructions to maximize accuracy and performance.
  • Monitor AI costs closely, as cheaper models can provide significant margin improvements without sacrificing accuracy.
  • Implement robust verification processes to ensure AI outputs are correct, especially for critical tasks.
Data Points
  • 91.0% (The highest accuracy achieved by the best-tuned AI model in Rippling's study.)
  • 88.5% to 89.5% (The accuracy range for seven AI models tested by Rippling.)
  • $621 (The cost of GLM 5.2, which achieved 88.7% accuracy without tuning.)
  • $2,509 (The cost of Opus 5, which performed similarly to Grok 4.5 but was significantly more expensive.)
  • 71 seconds to 131 seconds (The increase in response time when comparing Grok 4.5 to Grok 4.6.)

RevBots.ai View:

AI model performance benchmarks in production systems reveal cost and accuracy trade-offs that matter for GTM efficiency.

Full Story: SaaStr →