preloader

1 Post

Tag: Quantization

post-thumb

BY ericauthor-thumb

Sep 19, 2026

Now It Remembers: A 27B Model With a 262K-Token Context on the Same RTX 3060

The fifth in our RTX 3060 series. The same Qwen3.8-27B we ran last month, re-quantized to under 6GB, now runs entirely on the 12GB card, about four times faster, and holds up to 262,000 tokens of context. A whole novel, on a five-year-old gaming card.

agico

We transform visions into reality. We specializes in crafting digital experiences that captivate, engage, and innovate. With a fusion of creativity and expertise, we bring your ideas to life, one pixel at a time. Let's build the future together.

Copyright ©  2026  TYO Lab · v0.0.28