The "output can be checked in seconds" filter is the most practical heuristic I've seen for this question.
I've covered dozens of enterprise AI implementations, and the pattern holds at scale too—not just small business. The tools that survive are the ones where verification cost is low relative to the value delivered. When trust becomes the bottleneck instead of capability, adoption stalls regardless of how impressive the underlying technology is.
Your point about fully autonomous agents overpromising matches what I'm documenting in agentic AI coverage right now. The gap isn't model capability—it's operational infrastructure. Context handoffs, exception handling, and accountability all break down without human checkpoints, which is exactly why "unsupervised business function" agents demo well and fail in production.
The integration seam observation is the sharpest insight here. I see this constantly in larger organizations too: individual AI tools work fine in isolation, but nobody owns the handoff between them. That's where the real friction lives, and it's underpriced because it's not a flashy feature—it's plumbing.
Good, grounded framework. The three-bucket breakdown (generation, interaction, analysis) is a useful filter for anyone evaluating tools rather than chasing capability demos.