AMD posted results of a new study of the performance of the Ryzen AI Halo platform, stating that in the era of autonomous AI agents, the determining factor is not the speed of GPU token generation, but the speed of completing the entire workflow.
Unlike traditional chatbots, modern AI agents perform multi-stage tasks — they analyze documents, perform OCR, build a knowledge base, perform a search, verify the results, and only then generate a response.
That is why AMD suggests evaluating performance not by Tokens per Second or Time to First Token, but by much more important business indicators:
- End-to-End Completion Time;
- Cost per Completed Workflow;
- agent cycle execution speed;
- local inference performance.

Transition from chatbots to the Agentic Era
HEPA Benchmark Results
| Comparison metric | AMD Ryzen AI Max+ 395 (Halo) | NVIDIA DGX Spark (GB10) | AMD advantage |
| CPU Orchestration (Cleaning/OCR/Embedding) | 152.1 sec | 229.9 sec | 🚀 34% faster (71.5s vs. 139.6s on embeds) |
| Total execution time (End-to-End) | 311.6 sec | 367.1 sec | ⏱️ 15% faster (55.5 seconds ahead) |
| Cost per task (3-year CapEx) | ~ $ 0.0132 | ~ $ 0.0182 | 💰 27% cheaper than a processed workflow |
┌─ 367.1s │ │workflow executed│ │AMD 34% faster │ │55.5s saved │ │Combined RAM │ └─
HiTech Expert Take
AMD’s research demonstrates a significant shift in the entire generative AI industry. While in 2023–2025, manufacturers competed primarily on the speed of GPU token generation, in 2026 the focus shifts to the performance of complex agent systems. For corporate users implementing on-premises AI solutions, the time to complete a complete business process, data protection, and the cost of a single completed task become critical.
This trend seems especially relevant against the backdrop of the rapid growth of investments in AI infrastructure. TrendForce forecast, the combined capital expenditure of the world’s nine largest cloud providers will exceed $886 billion in 2026, indicating unprecedented demand for AI computing. At the same time, more and more companies are looking to local inference as a way to reduce their reliance on cloud APIs, reduce operational costs, and improve information security.
For Ukraine, the development of AI PC class and local agent systems opens up prospects for using generative AI in the public sector, industry, finance, medicine and defense technologies without the need to transfer confidential data to external services. In combination with modern cloud infrastructure, data centers and high-speed networks, such solutions can become an important element digital business transformation and GovTech.










