Apple has recently overhauled its entire M-Series chip plans, scrapping the launch of the M6 Pro and M6 Max processors and jumping straight to the M7 series. While the base M6 SoC is expected to launch, Apple is moving into the M7 series without the M6 Pro/Max variants, with plans to offer some...
Local LLMs are cool but also pretty slow compared to cloud. If you have to wait half an hour for your Feature while coding you might still opt for the cloud agent.
Have you tried running a local model on a M series Mac?
Yes ofc I ran Gemma 4 for example, but compare that to the speed of Gemini in the cloud the difference is massive.
How much RAM do you have and which version of the model did you run?
Local LLMs can be just as fast as long your device clears the requirements. If you noticed a huge difference, there’s a really good chance that you tried to use a model that requires more RAM than you have
Actually they can be much faster given sufficient VRAM and not a lot of concurrent users.
Yes, they are slower. However, I think that the pricing we’re going to see from the cloud providers might be enough to deter quite a lot of people. At least I hope so:
The fact that we’re already used to blazing speed generation kinda sucks. Local models are a much more sustainable way of unlocking the benefits of LLMs than giant ecosystem- and community-destroying data centers.