{"id":10361,"date":"2019-01-28T08:52:01","date_gmt":"2019-01-28T08:52:01","guid":{"rendered":"https:\/\/www.techdesignforums.com\/practice\/?p=10361"},"modified":"2019-01-28T08:55:40","modified_gmt":"2019-01-28T08:55:40","slug":"emulation-for-ai-part-two","status":"publish","type":"post","link":"https:\/\/www.techdesignforums.com\/practice\/technique\/emulation-for-ai-part-two\/","title":{"rendered":"Emulation for AI: Part Two"},"content":{"rendered":"<p><em><a href=\"https:\/\/www.techdesignforums.com\/practicetechnique\/emulation-for-ai-part-one\/\">The first part of this feature<\/a> discussed the trends causing AI companies to move from generic processors to their own ASICs or toward the use of still more focused platforms. This second article looks at what led Wave Computing, a pioneer in dedicated AI silicon \u2013 to adopt emulation during the development of its \u2018dataflow processing unit\u2019.<\/em><\/p>\n<p>Scalability is key to artificial intelligence (AI). According to open.ai, deep neural networks (DNNs) are doubling their performance demands every three-and-a-half months, compared to the 18 months of Moore\u2019s Law.<\/p>\n<p>At the same time, some applications are already beginning to strain the GPU-based AI processors that dominate AI silicon today. An analysis by Moore Insights &amp; Strategy (MI&amp;S) cites the examples of reinforcement learning for visual recognition and natural language processing as fields \u201cwhere accelerators sometimes struggle with more than a few nodes\u201d. Fragmentation in AI use-cases is thought likely to provide further examples.<\/p>\n<h3><strong>Wave Computing and the AI dataflow processing unit<\/strong><\/h3>\n<p><a href=\"https:\/\/wavecomp.ai\/\">Wave Computing<\/a> is a high profile Silicon Valley AI startup. It recently closed an $86M venture capital funding round and has open-sourced the MIPS processor instruction set architecture, which it acquired in 2018. Its main IP, however, is a dataflow processing unit (DPU), the basis for chips which it is now offering under early access for AI applications at the server, enterprise and edge levels (a further reflection of how AI processing is itself being distributed as applications proliferate).<\/p>\n<p>The DPU has been designed to harmonize with the dataflow graph of a DNN. In a paper commissioned by Wave, MI&amp;S describes the core processor architecture:<\/p>\n<p>\u201cThe company\u2019s implementation of a dataflow architecture seems to present an elegant alternative to train and process DNNs for AI, especially when models require a high degree of scaling across multiple processing nodes. Instead of building fast parallel processors to act as an offload math acceleration engine for CPUs, this dataflow machine directly processes the flow of data of the DNN itself. The CPU is only used as a pre-processor to initiate runtime, not as a workload scheduler, parameter server, or code\/data organizer.\u201d<\/p>\n<p>To address scalability, Wave has developed a technique that combines data parallelism and model parallelism. Data parallelism applies one model to every thread with different parts of the input spread across the processing resource. Model parallelism receives the same data but with the model split across the resource. The perceived benefits of model parallelism come when the model cannot be accommodated by a single acclerator.<\/p>\n<p>Then, to allow its silicon to be stacked, Wave has used a distributed signal manager to deliver data more efficiently to each processing unit and a combination of novel interconnect (which enables the signal management) and memory. M&amp;IS says of this:<\/p>\n<p>\u201cThe DPUs are interconnected directly with each other over a fabric (used for signaling \u201cFire\u201d and \u201cDone\u201d) and through dual ported Hybrid Memory Cubes (HMC), which act both as fast memory and as shared data buffers between the DPUs. This allows shared double buffering to improve scalability by keeping the critical data close to the processors.<\/p>\n<p>\u201cThe bisection bandwidth of this approach is impressive, which supports [Wave\u2019s] scale-up and scale-out thesis, delivering up to 7.25TB\/s per second. The company suggests that this would translate to approximately 300GB\/s of user data on average, since much of the bandwidth is consumed by the movement of feature map data moving between the agents in the DPUs.\u201d<\/p>\n<p>A block diagram for a four-CPU board is shown in Figure 1. <a href=\"https:\/\/wavecomp.ai\/research-paper-designed-to-scale\">M&amp;IS has produced a white paper that takes a more detailed look at Wave&#8217;s architecture<\/a>.<\/p>\n<div id=\"attachment_10362\" style=\"width: 1004px\" class=\"wp-caption aligncenter\"><a href=\"https:\/\/www.techdesignforums.com\/practicefiles\/2019\/01\/Wave_Architecture_Emulation_AI.jpg\"><img loading=\"lazy\" decoding=\"async\" aria-describedby=\"caption-attachment-10362\" class=\"size-full wp-image-10362\" src=\"https:\/\/www.techdesignforums.com\/practicefiles\/2019\/01\/Wave_Architecture_Emulation_AI.jpg\" alt=\"Figure 1. How AI dataflow processing units are interconnected through a bespoke fabric (Wave Computing)\" width=\"994\" height=\"587\" srcset=\"https:\/\/www.techdesignforums.com\/practice\/files\/2019\/01\/Wave_Architecture_Emulation_AI.jpg 994w, https:\/\/www.techdesignforums.com\/practice\/files\/2019\/01\/Wave_Architecture_Emulation_AI-300x177.jpg 300w, https:\/\/www.techdesignforums.com\/practice\/files\/2019\/01\/Wave_Architecture_Emulation_AI-768x454.jpg 768w, https:\/\/www.techdesignforums.com\/practice\/files\/2019\/01\/Wave_Architecture_Emulation_AI-650x384.jpg 650w\" sizes=\"auto, (max-width: 994px) 100vw, 994px\" \/><\/a><p id=\"caption-attachment-10362\" class=\"wp-caption-text\">Figure 1. How AI dataflow processing units are interconnected through a bespoke fabric (Wave Computing)<\/p><\/div>\n<h3><strong>Wave Computing and emulation<\/strong><\/h3>\n<p>\u201cThese are very big chips,\u201d says Jean-Marie Brunet, marketing director of Mentor\u2019s emulation business. \u201cThere are billions of gates \u2013 we\u2019re typically looking at three-to-four billion gates right now for AI &#8211; and according to its specification there are about 16,000 processing units on the main Wave chip.\u201d<\/p>\n<p>Wave shifted from an FPGA-based development strategy to emulation \u2013 and specifically to <a href=\"https:\/\/www.mentor.com\/products\/fv\/emulation-systems\/?cmpid=10166\">Mentor\u2019s Veloce emulation platform<\/a> \u2013 largely because of the scalability required by that silicon size.<\/p>\n<p>\u201cWith the demands on AI now, there are two ways you can scale: You put more hardware into the stack or you increase the capacity of the chip,\u201d says Brunet, \u201cand those both create capacity issues for FPGA.<\/p>\n<p>\u201cBut another aspect for Wave was that before they decided to really push on these DPUs, the software they were running on FPGA prototypes was end-user software, not performance software. When you move into your own performance software, you need a platform on which to run the right AI benchmarks \u2013 MLPerf and others. So, you now need many more metrics \u2013 power and performance &#8211; and visibility.<\/p>\n<p>\u201cThen, so much of this was new \u2013 there was no legacy \u2013 that for both the software and the hardware there was a real need for virtualization. You don\u2019t have a previous design you can just plug into ICE [in-circuit emulation]. It\u2019s all about models. The other issue with virtualization is that you can give highly distributed teams access to the platform and ensure concurrency of the hardware and software. [Wave has offices in the US, China, Taiwan, Sri Lanka and the Philippines]\u201d<\/p>\n<p>Scalability and virtualization form two of what Brunet sees as the three \u2018pillars\u2019 supporting the case for emulation in AI, with determinism as the third.<br \/>\n\u201cAnd again, that is something that made Wave a good example \u2013 once they moved from end-user software to their own, they needed a platform to which the design would map exactly in the same way for each iteration. If you don\u2019t have that and you\u2019re looking to made architectural tradeoffs, then you\u2019re wasting your time.\u201d<\/p>\n<h3><strong>A building emulation market in AI<\/strong><\/h3>\n<p>If one takes Wave as emblematic of a trend, then things look good for emulation in the AI business. According to AI analyst Cognilytica, venture capital funding for AI startups reached $17B in 2017 and almost certainly exceeded that figure in 2018. The money is out there to add silicon NRE to AI\u2019s traditional foundations in algorithms and software.<\/p>\n<p>And the traditional silicon and system companies are getting involved. Companies like Google and Facebook have in-house projects, Nvidia is heavily committed (and itself responding to both the scalability and fragmentation issues) as is Intel. Qualcomm set up a $100M corporate venturing fund targeting AI at the end of 2018, and Wave itself has partnered with Broadcom for the development of its 7nm generation.<\/p>\n<p>\u201cAnd there are no real standards, so it is a bit like the Wild West, and that is obviously good for us,\u201d says Brunet. \u201cI think there are a couple of analogies you can look at.<\/p>\n<p>\u201cIf you go back about 25 years ago and look at networking, it started that software was the key thing, and you put a few CPUs in a workstation. But then people realized that the power and the performance weren\u2019t there and they would have to do their own hardware designs. It was expensive to start with, but soon the economies of scale kicked in. It\u2019s much the same in AI now.<\/p>\n<p>\u201cThe other more recent one to look at is what happened when we didn\u2019t have EUV and we started with double-patterning, then multi-patterning (double, triple, quadruple). There was no single approach. Like with AI, we had to support them all.\u201d<\/p>\n<p>There is the issue that these will often be originally software led-companies coming to AI, but here Brunet is phlegmatic.<\/p>\n<p>\u201cIf you start with the big systems companies, then obviously they already have silicon expertise in-house, so they know about emulation. They know about doing very big, very complex chips and what emulation can do,\u201d he says. \u201cThe startups are having to bring in engineering expertise, who also bring an understanding of emulation with them..<\/p>\n<p>\u201cThe basic hardware arguably isn\u2019t as complex as some other designs. But these are big chips and they have to scale. If you look at the trends: If a company wins a socket they know that the AI or ML or DL system is going to scale and scale rapidly, so they have to be able to scale with it.<\/p>\n<p>\u201cThere is an incredibly strong argument that emulation has a place here.\u201d<\/p>\n","protected":false},"excerpt":{"rendered":"<p>The second part of this feature looks at how Wave Computing&#8217;s objectives with its dataflow processing unit for AI mapped to the use of emulation in its development.<\/p>\n","protected":false},"author":138,"featured_media":10358,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"footnotes":""},"categories":[38],"tags":[2045,2276,871,2159,2024,2280,1903,1047,1257,1532,2281,2016,2272,1741,1017],"coauthors":[1441],"class_list":["post-10361","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-ic-implementation","tag-ai","tag-algorithm","tag-architecture","tag-artificial-intelligence","tag-caffe","tag-dataflow-processing-unit","tag-deep-learning","tag-emulation","tag-gpu","tag-hardware-software-co-design","tag-interconnect","tag-machine-learning","tag-mlperf","tag-parallelism","tag-validation","workflow-analysis","workflow-expert-blog","workflow-interview","workflow-up-to-date","organization-caffe","organization-mentor","organization-tensorflow","organization-wave-computing"],"_links":{"self":[{"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/posts\/10361","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/users\/138"}],"replies":[{"embeddable":true,"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/comments?post=10361"}],"version-history":[{"count":0,"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/posts\/10361\/revisions"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/media\/10358"}],"wp:attachment":[{"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/media?parent=10361"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/categories?post=10361"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/tags?post=10361"},{"taxonomy":"author","embeddable":true,"href":"https:\/\/www.techdesignforums.com\/practice\/wp-json\/wp\/v2\/coauthors?post=10361"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}