Different model choices affect cost, speed, quality, reliability, workflow and scalability.
I do not think we have a final model-selection framework. But we have observed enough to know that treating “AI” as one interchangeable capability is too simple.
Cost made the problem visible
We observed model usage increasing significantly without a corresponding increase in useful work completed. That forced a practical question:
Are we using more model capability than the task actually requires?
Quality is not one thing
The useful question became less “Which model is best?” and more “What does this task actually require?”
Current working distinction
Low-complexity tasks may include formatting, metadata, simple extraction and repetitive structured work.
Medium-complexity tasks may include account analysis, comparing known criteria and drafting from structured context.
High-complexity tasks may include architecture, difficult debugging, strategic synthesis, new method development and ambiguous evidence.
This is a practical heuristic, not a validated taxonomy.
Some things we currently think we have observed
- Different AI tasks appear to need meaningfully different levels of model capability.
- Higher capability can increase cost much faster than it increases value on simple tasks.
- Some difficult tasks appear to benefit substantially from stronger reasoning.
- Workflow design can sometimes reduce the need for a stronger model.
- Verification difficulty seems important when choosing how much model capability to use.
- The economics of model choice become more important as usage scales.
What we do not know yet
How should task complexity be measured? When should escalation happen automatically? How much quality improvement justifies higher cost? How stable are routing rules as models improve?
Try the idea
A later interaction can give the visitor three tasks — extraction, commercial analysis and architecture review — and ask which level of model capability they would choose and why.