INTEL DOSSIER|CLASSIFIED
DECRYPTEDFile #16 • October 1, 2025
Building AI Learning Tools: Our Mission
projectsai-toolstesting
Authorcoderunner
Categoryprojects
StatusPUBLISHED
ClearancePUBLIC
//This blog has covered AI tools, models, providers, and optimization. Now we're building our own tools to systematically test and compare AI capabilities. This post is our public roadmap.
We’ve spent months analyzing AI tools from the outside. Now we’re building our own.
The Problem
Most AI benchmarks are:
- Synthetic: Test datasets that don’t reflect real work
- Opaque: Black-box scores without methodology
- Static: One-time results that go stale quickly
- Incomplete: Test one dimension while ignoring others
We need systematic, repeatable, real-world testing.
What We’re Building
1. AI Capability Test Suite
A standardized test suite covering:
- Reasoning: Logic puzzles, multi-step problems, edge cases
- Code: Generation, refactoring, debugging, architecture
- Writing: Technical docs, creative, editing, style consistency
- Math: Arithmetic, algebra, calculus, proofs
- Safety: Jailbreak resistance, hallucination detection, bias
- Speed: TTFT, TPS, total latency under load
- Cost: Per-task cost at various quality levels
2. Model Comparison Dashboard
Real-time comparison of models across dimensions:
- Quality scores by category
- Cost per task
- Speed metrics
- Value ratio (quality per dollar)
- Trend lines over time
3. Harness Testing Framework
Test how models behave under different conditions:
- Prompt engineering techniques
- RAG with various chunk sizes
- Fine-tuning impact
- Temperature and parameter sensitivity
- Multi-turn conversation quality
4. Automated Regression Testing
For production AI applications:
- Track model updates for quality changes
- Alert when quality drops below thresholds
- A/B test prompt changes
- Cost tracking and optimization alerts
Development Roadmap
Phase 1: Foundation (Q4 2025)
- Core test suite with 500+ test cases
- Support for 10+ models via OpenRouter
- Basic reporting and comparison
- Public GitHub repo
Phase 2: Automation (Q1 2026)
- Automated daily benchmark runs
- Historical trending dashboard
- API for programmatic access
- CI/CD integration
Phase 3: Community (Q2 2026)
- Community submissions for test cases
- Shared benchmark results
- Leaderboard across models
- Plugin system for custom tests
Phase 4: Products (Q3 2026+)
- Hosted benchmark service
- Production monitoring tools
- Optimization recommendations
- Enterprise features
Current Progress
Completed
- Hermes Agent deep dive (tool, plugins, skills, providers)
- Claude Code, Cursor, Codex, OpenCode/KiloCode reviews
- LLM provider landscape analysis
- OpenRouter model benchmarks (8 models, 600 tests)
- Output improvement methodology
In Progress
- AI Harness testing framework
- Systematic capability test suite
- Model comparison dashboard
Planned
- Automated daily benchmarks
- Cost optimization tools
- Production monitoring
How to Follow Along
- Blog posts: Every tool and model we test gets a detailed review
- Open source: All tools will be on GitHub
- Benchmark data: Raw data published for every test
- Newsletter: Monthly summaries of findings
Get Involved
We’re building in public. Contributions welcome:
- Test cases: Submit real-world tasks to benchmark
- Models: Request specific models or providers to test
- Tools: Suggest tools to review or build together
- Data: Share your own benchmark results
The goal is simple: build the definitive resource for AI model capabilities, tested on real work, shared with everyone.
Detailed Analysis
Strengths4 PROS
- +Hands-on experience with every model and provider
- +Real benchmarks on real tasks, not synthetic tests
- +Open source tools the community can use
- +Building expertise to make better recommendations
Weaknesses4 CONS
- −API costs for comprehensive testing add up
- −Tool development takes time away from content
- −Results may be biased by our specific use cases
- −Maintenance burden for open source projects
PricingOpen Source
All tools we build will be open source. Costs are infrastructure and API credits for testing.
END OF FILE|DISTRIBUTION: UNLIMITED
← Back to Archive