A hardware enthusiast acquired an Nvidia V100 16GB GPU for ~$100 due to its SXM2 socket (server bus, not PCIe), then used a ~$100 adapter board to connect it to a consumer motherboard. Despite being a 2017 card, it outperforms an RTX 3060 12GB in tokens-per-second for local LLM inference, though at higher idle power. The arbitrage window may close as the market catches on, but for now it's a cost-effective path to running large open-source AI models locally.
165 Impressions