Qwen 3.8 27B: 16GB VRAM Local Test
A hands-on test of the newly released Qwen 3.8 27B model running locally on a 16GB VRAM RTX 5060Ti via llama.cpp. The dense hybrid model (linear attention on 48 of 64 layers plus gated attention) supports native multimodal input including video, a 262K context window expandable to 1M tokens, and toggleable thinking mode. Coding tasks like a portfolio site and a 3D racing game were generated but were slow (4.78-6.42 tokens/sec) and required long generation times (over an hour). Benchmark scores are compared against Opus 4.6 Max, showing gains on terminal-bench, LiveCodeBench, DeepSeek 1.1, and OSWorld Verified.