GLM 5.2 4bit (500GB total weights) runs at 35 t/s generation, 2000 t/s prefill on DGX Station with DwarfStar mixed RAM/VRAM inference, soon likely to reach 3k t/s.
GLM 5.2 4bit Achieves 35 T/s Generation on DGX Station with DwarfStarAI News
GLM 5.2 4bit (500GB total weights) runs at 35 t/s generation, 2000 t/s prefill on DGX Station with DwarfStar mixed RAM/VRAM inference, soon likely to reach 3k t/s.
