I subscribe to the idea that the not too distant future there will just be a compute "fabric" of every device mobile to rooms of server racks will be running the model, or whatever we call it after that. Projects like https://github.com/Mesh-LLM/mesh-llm are a glimpse. All the pieces are maturing in own silos until put together well.