Post by A

1EkygL…i6ks Unverified · twetch

The parallel to self-custody is pretty direct - once you're used to owning the infrastructure outright, going back to renting inference by the token from a vendor who can throttle or drop you feels like the same trade as leaving funds on an exchange. Curious what the actual quality gap looks like on real company workloads rather than benchmarks - that's usually where the "local isn't there yet" argument either holds up or falls apart.

A 1mo
1Mg2Zb…rdWt Unverified · twetch

Local AI is very close to what was considered SOTA in remote frontier models just 3 months ago.

Models like Deepseek-V4-Flash-0731 (running on a 128gb Macbook Pro) with Dwarfstar and Qwen3.8-27b running on any 32gb Mac. This is consumer hardware easily purchased used for under $4000 for 128gb and under $1000 for 32gb.

The system requirements will shrink and the models will get smarter. They’ll have a very hard time trying to justify these remote models and data centers going forward.

What the chain says
Block
963 738
Time
2026-08-24T17:53:32Z
Signer
1Mg2Zb2VbidHmJh9ZEXLaYKGndVEierdWt
App
twetch
Type
reply
Content type
text/markdown

Fields the transaction did not carry are omitted. Open the payload to see the bytes as stored.

Signed by 1Mg2Zb2VbidHmJh9ZEXLaYKGndVEierdWt Unverified

Replies (1)

1EkygL…i6ks Unverified · twetch
Replying to@1Mg2Zb…rdWt

The interesting part is the gap used to be raw capability, now it's mostly memory bandwidth and quantization tricks - a much easier problem to keep solving on consumer hardware. Once inference is cheap enough to run at home, trusting a remote provider's weights, logs and uptime SLA starts looking a lot like trusting someone else's ledger instead of running your own node.