Ссылка
click to show
click to show
Tiiny AI Pocket Lab a 65W 190TOPS 80GB poket PC.
Summary
Tiiny Ai releases a mini PC with a dNPU and 80GB of LPDDR5X
Quotes
Quote
Key Specifications:
Processor: ARMv9.2 12-core CPU
AI Compute Power: Custom heterogeneous module (SoC + dNPU), delivering ~190 TOPS
Memory & Storage: 80 GB LPDDR5X + 1 TB SSD
Model Capacity: Runs up to 120B-parameter LLMs fully on-device
Power Efficiency: 30 W TDP, 65 W typical system power
Dimensions & Weight: 14.2 × 8 × 2.53 cm, approx. 300 g, pocket-sized
Ecosystem: One-click deployment of dozens of open-source LLMs and agent frameworks
Connectivity: Works fully offline; no internet or cloud required
The Tiiny AI Pocket Lab will be launched on Kickstarter within the next few months and will retail for $1,399 – pricey for a mini-computer,
My thoughts
The device is interesting, it's like a NAS for AI or like an eGPU. You connect to it, and it provides I suspect a web interface to use the models.
The specs are deficient in some key specs. Memory speed and bus width (I suspect dual channel 128b perhaps 8000). Quantization supported, likely some form of 4 bit quants. Who provides the dNPU.
At first glance it competes with the "AI mini supercomputers" like the AMD AI MAX and Nvidia DGX Spark, but not really.
Nvidia and AMD models use regular shaders (and to a smaller extent some special tensor acceleration that is really hard to use), will run standard pytorch and llama.cpp application as long as the vendor provides the stack for it you can run all the models.
NPU are theoretically much more area efficient and much much more energy efficient at doing tensor operation than regular shaders, for ML workloads. The drawback is you need special software to run acceleration and special software to convert model into the supported format, meaning the NPU provider has to do lots of continuous software work to maintain that piece of hardware. Training is likely a big no.
They cite Qwen 3 and Zimage in the compatibility that happens to be popular and powerful right now.
Tiling compute units is the fun part, doing a stack that keeps them fed is why Nvidia is worth at last a few trillions. E.g. For my 7640u AMD stil doesn't support 760m ROCm acceleration, and I don't think anything runs on its NPU at all.
The device is really interesting, promising to be very efficient for inference on the go and not having to deal with the ML dependencies on the main machine. On the flip side, software support is likely going to be extremely hard to maintain. New models keep being released faster, and they need to be ported to this NPU. If the development behind it lapse just