AMD Ryzen AI Max+ 395 Local AI Workstation, 128GB — 96GB VRAM | Alden Tech
Local AI, Low Power
$3,799
Built around AMD's Ryzen AI Max+ 395 platform with 128GB of high-speed unified memory instead of 64GB, with up to 96GB configurable as dedicated GPU memory (VGM) — and per AMD's own spec, total graphics-addressable memory can reach about 112GB on this configuration, since the GPU can also draw on shared memory beyond the VGM figure depending on workload and backend. That headroom fits significantly larger local models than the 64GB configuration can run at all: practical territory includes 70B, 100B+, and large mixture-of-experts models, depending heavily on quantization and context length, plus better concurrency and larger RAG context windows. The Radeon 8060S integrated GPU runs on AMD's ROCm 7 platform, backed by an NPU rated up to 50 TOPS (126 TOPS combined system-wide). Configured with your choice of Ubuntu Desktop 24.04 or Windows 11 Pro — Linux has the broader ROCm-based serving ecosystem on this platform, but Windows is a fully viable path too; the real difference is tool-specific, not a blanket OS ranking. Ships with LM Studio installed and configured by default — the same Vulkan/llama.cpp backend on either OS. Ollama and standalone llama.cpp are available through the AI Tooling Package for more advanced setups: Ollama runs on Vulkan on Windows (its Linux ROCm support doesn't extend to Windows) or natively via ROCm 7 on Ubuntu; llama.cpp supports either Vulkan or ROCm/HIP on both operating systems.
AI Profile
- Total VRAM: 96GB
- GPU: 1x AMD Radeon 8060S Graphics
Workload Categories
- Local LLM inference
- Private company AI assistant
- RAG / document search
- AI-assisted development
- Local image generation
- AI experimentation
- Multi-user internal AI
- Sensitive-data workflows
Specifications
| cpu | AMD Ryzen™ AI Max+ 395 |
|---|---|
| gpu | AMD Radeon 8060S Graphics |
| AI Engine Performance | AMD Ryzen™ AI NPU Computing Power: Up to 50 TOPS | Total Computing Power: Up to 126 TOPS |
| memoryCapacity | 128GB LPDDR5x-8000MT/s |
| storage | 2TB PCIe 4.0 |
| cooling | Copper Base, 6 Heat Pipes Dual Fans & Phase-Change Cooling 160W Peak, 130W Sustained |
| networking | 2× 10GbE LAN (RJ45) - Wi-Fi 7 | Bluetooth® 5.4 |
| operatingSystem | Ubuntu Desktop 24.04 or Windows 11 Pro |
| powerSupply | AC INPUT (100-240V ~6A 50-60Hz, Internal DC 12V/26.6A, 320W MAX) |
| LLM Software & Tooling | LM Studio included by default (same Vulkan/llama.cpp backend on Windows or Ubuntu). Ollama and llama.cpp also available via the AI Tooling Package — support is tool-specific, not an OS-blanket rule: Ollama is Vulkan-only on Windows (its Linux ROCm support doesn't extend to Windows) and runs natively via ROCm on Ubuntu; llama.cpp supports either Vulkan or ROCm/HIP on both operating systems. |
| Practical Model Range | ~70B–100B+ and large MoE models (quantization/context dependent) |
| System Access | Full SSH and root/administrator access — same as any PC we sell. Nothing is locked down or requires our involvement. |
Available Options
- Preinstalled AI Tooling Package (+$250) — Ollama and llama.cpp (alternative inference backends), Open WebUI, and RAG document ingestion pipeline — installed and configured before pickup.
There is no online checkout for this system yet — mention any option you want when you contact us.
Warranty
3-Year Parts & Labor Warranty — see the full policy.
LLM Benchmark Results
These are published third-party benchmark results for the AMD Ryzen AI Max+ 395 platform this system is built on — not tests performed by Alden Tech, and not independently reproduced by us. Results below used different operating systems, backends (ROCm vs. Vulkan), model files, quantizations, prompt lengths, and context settings, so they are not directly comparable to each other. A configured context limit is not the same as a demonstrated occupied prompt length — see each result's notes. Sources are linked on each result.
| Model | Prompt tok/s | Generation tok/s | Backend | OS | Notes |
|---|---|---|---|---|---|
| LFM2-24B-A2B (Q4_K_M) | 1190 | 109 | Vulkan | — | 7,624-token document summarization task. Author reported minor summary-quality issues. One task, not an explicit concurrent serving test. [BDOC-01] Source |
| GPT-OSS 20B (Q8_0) | 1040 | 70 | Vulkan | — | 7,624-token document summarization task. Author reported a favorable speed/quality combination. [BDOC-02] Source |
| Qwen3-Coder-30B-A3B (Q4_K_M) | 779 | 66 | Vulkan | — | 7,624-token document summarization task. Coding-oriented MoE evaluated on a summarization workload, not its primary use case. [BDOC-03] Source |
| Nemotron-3-Nano-30B (Q4_K_M) | 798 | 61 | Vulkan | — | 7,624-token document summarization task. [BDOC-04] Source |
| GPT-OSS 120B (Q4_K_M) | 404 | 52 | Vulkan | — | 7,624-token document summarization task. Large MoE result for this specific summary workload — not proof every 120B model or context length performs similarly. [BDOC-05] Source |
| Qwen3.5-35B-A3B UD (Q4_K_M) | 691 | 49 | Vulkan | — | 7,624-token document summarization task. UD variant; a reasonable general business-assistant candidate, though this is not a standardized quality evaluation. [BDOC-06] Source |
| GLM-4.7-Flash (Q5_K_M) | 484 | 48 | Vulkan | — | 7,624-token document summarization task. Author reported a formatting issue on this prompt; throughput does not establish output quality. [BDOC-07] Source |
| Devstral Small 2 (Q4_K_M) | 281 | 14 | Vulkan | — | 7,624-token document summarization task. Dense 24B model — decodes markedly slower than the MoE models above at similar parameter counts, illustrating why total parameters alone don't predict speed. [BDOC-08] Source |
| Qwen3.8-27B (UD-Q5_K_XL) | 31.44 | 16.45 | Vulkan + MTP | Windows 11 | Configured context 131,072 tokens — a --ctx-size setting, not a demonstrated fully-occupied prompt. Sustained 512-token reasoning benchmark; full GPU offload, q8_0 K/V cache, one slot, MTP 4 draft tokens, p_min 0.75. [BQ38-01] Source |
| Qwen3.8-27B (UD-Q5_K_XL) | 33.87 | 16.68 | Vulkan + MTP | Windows 11 | Configured context 65,536 tokens — not a demonstrated fully-occupied prompt. Same MTP/offload setup as the 131K run. Author's practical default is Q5 at 128K context, not this 64K configuration. [BQ38-02] Source |
| Qwen3.8-27B (Q6_K) | 28.93 | 14.24 | Vulkan + MTP | Windows 11 | Configured context 65,536 tokens — not a demonstrated fully-occupied prompt. Decoded more slowly than Q5 and used more memory in this run; no standardized quality comparison supplied. [BQ38-03] Source |
| Qwen3.8-27B (Q8_0) | 31.44 | 12.74 | Vulkan + MTP | Windows 11 | Configured context 65,536 tokens — not a demonstrated fully-occupied prompt. Highest-precision quant tested here; used the most memory and decoded slowest of the four. [BQ38-04] Source |
| Qwen3 Coder 30B (Q4_K_M) | 812 | 69.6 | ROCm 7.2 | Ubuntu 25.10 | Tested on Ubuntu 25.10 (kernel 6.17) — not the 24.04 this unit ships with; shown as representative Strix Halo/ROCm data. llama-bench pp512/tg128, three averaged repetitions, Flash Attention — a standalone microbenchmark, not a combined occupied-context or multi-user test. [BROC-01] Source |
| Llama 3.2 3B (Q6_K) | 1538 | 69 | ROCm 7.2 | Ubuntu 25.10 | Tested on Ubuntu 25.10, not the 24.04 this unit ships with. Small dense model; same pp512/tg128 microbenchmark methodology as the other BROC rows. [BROC-02] Source |
| DeepSeek Coder V2 (Q5_K_M) | 1227 | 68 | ROCm 7.2 | Ubuntu 25.10 | Tested on Ubuntu 25.10, not the 24.04 this unit ships with. Same pp512/tg128 microbenchmark, three repetitions. [BROC-03] Source |
| GLM-4.7 Flash (Q4_K_M) | 897 | 54.1 | ROCm 7.2 | Ubuntu 25.10 | Tested on Ubuntu 25.10, not the 24.04 this unit ships with. Same pp512/tg128 microbenchmark, three repetitions. [BROC-04] Source |
| GPT-OSS 120B (MXFP4) | 174 | 51.1 | ROCm 7.2 | Ubuntu 25.10 | Tested on Ubuntu 25.10, not the 24.04 this unit ships with. Same pp512/tg128 microbenchmark methodology. [BROC-05] Source |
| Llama 4 Scout (Q4_K_M) | 282 | 19.2 | ROCm 7.2 | Ubuntu 25.10 | Tested on Ubuntu 25.10, not the 24.04 this unit ships with. Same pp512/tg128 microbenchmark methodology. [BROC-06] Source |
| Qwen 3.5 122B-A10B (MXFP4) | 189 | 19.5 | ROCm 7.2 | Ubuntu 25.10 | SOURCE CONFLICT, unresolved: the project's README reports 189/19.5 (preserved here), but its own docs/MODELS.md reports 136/18.3 for the same test. Do not use either number in an unqualified headline claim until resolved. Tested on Ubuntu 25.10, not the 24.04 this unit ships with. [BROC-07] Source |
| Qwen3-235B Thinking (Q3_K_M) | 135 | 7.8 | ROCm 7.2 | Ubuntu 25.10 | Tested on Ubuntu 25.10, not the 24.04 this unit ships with. Partial GPU offload (80/95 layers) — this model loads but does not represent typical interactive-office speed at this size. Capacity-vs-speed example, not a recommended everyday configuration. [BROC-08] Source |
| Llama 3.3 70B (Q6_K) | 86 | 3.8 | ROCm 7.2 | Ubuntu 25.10 | Tested on Ubuntu 25.10, not the 24.04 this unit ships with. Dense 70B model — low decode speed illustrates why parameter count alone doesn't predict usability; a MoE of similar total size performs very differently (see Qwen 3.5 122B-A10B above). [BROC-09] Source |
| Qwen3.6-35B-A3B (Q4_K_M) | — | — | Vulkan | Linux | Derived pace ≈62.9 tok/s, calculated as 1000 / 15.9ms median inter-token latency — not a directly reported generation rate. Three repeated runs, one physical unit, 100% completion. [BCON-01] Source |
| Qwen3.6-35B-A3B (Q4_K_M) | — | — | Vulkan | Linux | 3-hour sustained 4-concurrent-request workload across two units (one 3-hour pass each). Derived pace ≈33.3 tok/s, calculated as 1000 / 30.0ms median inter-token latency — this is NOT an aggregate throughput and should not be multiplied by 4. 100% completion, 1.6% decode-pace drift over the 3 hours. [BCON-02] Source |
| Qwen3.6-35B-A3B (Q4_K_M) | — | — | Vulkan | Linux | 8 concurrent requests: 100% completion, 865ms time to first answer (not first token). No generation or prompt throughput rate was reported for this configuration. [BCON-03] Source |
See the full Business Hardware lineup: Business Hardware.
Sold by Alden Tech — veteran-owned custom PC builder in Spartanburg, South Carolina. Call or text (864) 381-8201. Contact