@nprofile1q... It's really not that difficult, and not that expensive. Most AI workloads are not that complex, and modern techniques to improve long-context performance make even on-device models capable of solving many rudimentary office tasks.
My consultancy does exactly this, there's at least 3 examples I can refer to with clients reducing their AI bills by 90% after switching to a Gemma or Qwen model running on a single RTX Pro 6000.
(Pictured is the performance improvement on a couple of long-context tasks, i.e. OOLONG, OOLONG-Pairs, and the step-change performance improvement after an 8B param model is given a Python REPL loop and ability to recursively call itself)
Source:
https://arxiv.org/abs/2512.24601