INTEL DOSSIER|CLASSIFIED
DECRYPTED
File #16 • October 1, 2025

Building AI Learning Tools: Our Mission

projectsai-toolstesting
Authorcoderunner
Categoryprojects
StatusPUBLISHED
ClearancePUBLIC
//This blog has covered AI tools, models, providers, and optimization. Now we're building our own tools to systematically test and compare AI capabilities. This post is our public roadmap.

We’ve spent months analyzing AI tools from the outside. Now we’re building our own.

The Problem

Most AI benchmarks are:

  • Synthetic: Test datasets that don’t reflect real work
  • Opaque: Black-box scores without methodology
  • Static: One-time results that go stale quickly
  • Incomplete: Test one dimension while ignoring others

We need systematic, repeatable, real-world testing.

What We’re Building

1. AI Capability Test Suite

A standardized test suite covering:

  • Reasoning: Logic puzzles, multi-step problems, edge cases
  • Code: Generation, refactoring, debugging, architecture
  • Writing: Technical docs, creative, editing, style consistency
  • Math: Arithmetic, algebra, calculus, proofs
  • Safety: Jailbreak resistance, hallucination detection, bias
  • Speed: TTFT, TPS, total latency under load
  • Cost: Per-task cost at various quality levels

2. Model Comparison Dashboard

Real-time comparison of models across dimensions:

  • Quality scores by category
  • Cost per task
  • Speed metrics
  • Value ratio (quality per dollar)
  • Trend lines over time

3. Harness Testing Framework

Test how models behave under different conditions:

  • Prompt engineering techniques
  • RAG with various chunk sizes
  • Fine-tuning impact
  • Temperature and parameter sensitivity
  • Multi-turn conversation quality

4. Automated Regression Testing

For production AI applications:

  • Track model updates for quality changes
  • Alert when quality drops below thresholds
  • A/B test prompt changes
  • Cost tracking and optimization alerts

Development Roadmap

Phase 1: Foundation (Q4 2025)

  • Core test suite with 500+ test cases
  • Support for 10+ models via OpenRouter
  • Basic reporting and comparison
  • Public GitHub repo

Phase 2: Automation (Q1 2026)

  • Automated daily benchmark runs
  • Historical trending dashboard
  • API for programmatic access
  • CI/CD integration

Phase 3: Community (Q2 2026)

  • Community submissions for test cases
  • Shared benchmark results
  • Leaderboard across models
  • Plugin system for custom tests

Phase 4: Products (Q3 2026+)

  • Hosted benchmark service
  • Production monitoring tools
  • Optimization recommendations
  • Enterprise features

Current Progress

Completed

  • Hermes Agent deep dive (tool, plugins, skills, providers)
  • Claude Code, Cursor, Codex, OpenCode/KiloCode reviews
  • LLM provider landscape analysis
  • OpenRouter model benchmarks (8 models, 600 tests)
  • Output improvement methodology

In Progress

  • AI Harness testing framework
  • Systematic capability test suite
  • Model comparison dashboard

Planned

  • Automated daily benchmarks
  • Cost optimization tools
  • Production monitoring

How to Follow Along

  • Blog posts: Every tool and model we test gets a detailed review
  • Open source: All tools will be on GitHub
  • Benchmark data: Raw data published for every test
  • Newsletter: Monthly summaries of findings

Get Involved

We’re building in public. Contributions welcome:

  • Test cases: Submit real-world tasks to benchmark
  • Models: Request specific models or providers to test
  • Tools: Suggest tools to review or build together
  • Data: Share your own benchmark results

The goal is simple: build the definitive resource for AI model capabilities, tested on real work, shared with everyone.

Detailed Analysis

Strengths4 PROS
  • +Hands-on experience with every model and provider
  • +Real benchmarks on real tasks, not synthetic tests
  • +Open source tools the community can use
  • +Building expertise to make better recommendations
Weaknesses4 CONS
  • −API costs for comprehensive testing add up
  • −Tool development takes time away from content
  • −Results may be biased by our specific use cases
  • −Maintenance burden for open source projects
PricingOpen Source

All tools we build will be open source. Costs are infrastructure and API credits for testing.

END OF FILE|DISTRIBUTION: UNLIMITED
← Back to Archive