{"id":8282,"date":"2024-05-28T07:40:23","date_gmt":"2024-05-28T07:40:23","guid":{"rendered":"https:\/\/www.fibermall.com\/blog\/?p=8282"},"modified":"2024-05-28T07:41:35","modified_gmt":"2024-05-28T07:41:35","slug":"how-does-dsa-outperform-nvidia-gpus","status":"publish","type":"post","link":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm","title":{"rendered":"How does DSA Outperform NVIDIA GPUs?"},"content":{"rendered":"\n<p>SambaNova Systems hailed as one of the top ten unicorn companies in the United States, raised $678 million in a Series D funding round led by SoftBank in April 2021, achieving a staggering $50 billion valuation. The company&#8217;s previous funding rounds involved prominent investors such as Google Ventures, Intel Capital, SK, Samsung Catalyst Fund, and other leading global venture capital firms. What disruptive technology has SambaNova developed to attract such widespread interest from top-tier investment institutions worldwide?<\/p>\n\n\n\n<p>According to SambaNova&#8217;s early marketing materials, the company has taken a different approach to challenge the AI giant NVIDIA.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambanova.png\" alt=\"sambanova\" class=\"wp-image-8284\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambanova.png 830w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambanova-300x149.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambanova-768x381.png 768w\" sizes=\"(max-width: 830px) 100vw, 830px\" \/><\/figure>\n\n\n\n<p>The comparison is quite astonishing: a single SambaNova machine is claimed to be equivalent to a 1024-node NVIDIA V100 cluster built with tremendous computational power on the NVIDIA platform. This first-generation product, based on the SN10 RDU, is a single 8-card machine.<\/p>\n\n\n\n<p>Some might argue that the comparison is unfair, as NVIDIA has the DGX A100. SambaNova seems to have acknowledged this, and their second-generation product, the SN30, presents a different comparison:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/DGX-A100.png\" alt=\"DGX A100\" class=\"wp-image-8285\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/DGX-A100.png 832w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/DGX-A100-300x143.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/DGX-A100-768x366.png 768w\" sizes=\"(max-width: 832px) 100vw, 832px\" \/><\/figure>\n\n\n\n<p>The DGX A100 offers 5 petaFLOPS of computing power, while SambaNova&#8217;s second-generation DataScale also delivers 5 petaFLOPS. In terms of memory, the DGX A100 has 320GB of HBM, while the SN30 boasts 8TB of DDR4 (the author speculates that this might be a typo, and the actual figure should be 3TB * 8).<\/p>\n\n\n\n<p>The second-generation chip is a Die-to-Die version of the SN10 RDU. The SN10 RDU&#8217;s architecture specifications are 320TFLOPS@BF16, 320M SRAM, and 1.5T DDR4. The SN30 RDU essentially doubles these specifications, as described below:<\/p>\n\n\n\n<p><em>\u201cThis chip had 640 pattern compute units with more than 320 teraflops of compute at BF16 floating point precision and also had 640 pattern memory units with 320 MB of on-chip SRAM and 150 TB\/sec of on-chip memory bandwidth. Each SN10 processor was also able to address 1.5 TB of DDR4 auxiliary memory.\u201d<\/em><em><\/em><\/p>\n\n\n\n<p><em>\u201cWith the Cardinal&nbsp;SN30 RDU, the capacity of the RDU is doubled, and the reason it is doubled is that SambaNova designed its architecture to make use of multi-die packaging from the get-go, and in this case SambaNova is doubling up the capacity of its DataScale machines by cramming two new RDU \u2013 what we surmise are two tweaked SN10s with microarchitectures changes to better support large foundation models \u2013 into a single complex called the SN30. Each socket in a DataScale system now has twice the compute capacity, twice the local memory capacity, and twice the memory bandwidth of the first generations of machines.\u201d<\/em><em><\/em><\/p>\n\n\n\n<p>High bandwidth and large capacity are mutually exclusive choices. NVIDIA opted for high bandwidth with HBM, while SambaNova chose large capacity with DDR4. From a performance perspective, SambaNova appears to outperform NVIDIA in AI workloads.<\/p>\n\n\n\n<p><strong>SambaNova SN10 RDU at Hot Chips 33<\/strong><strong><\/strong><\/p>\n\n\n\n<p>The Cardinal SN10 RDU is SambaNova&#8217;s first-generation chip, while the SN30 is the second-generation chip (a Die-to-Die version of the SN10). We notice its significant features, including a large on-chip SRAM, massive on-chip bandwidth, and an enormous 1.5T DRAM capacity.<\/p>\n\n\n\n<p>The SN40L is the third-generation chip (with added HBM). The first two generations of chips relied on the Spatial programming characteristics of Dataflow, reducing the demand for high DRAM bandwidth and opting for the large-capacity DDR route. However, the third-generation chip builds upon this foundation and incorporates 64GB of HBM, combining both bandwidth and capacity.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png\" alt=\"SambaNova SN10 RDU at Hot Chips 33\" class=\"wp-image-8286\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>This leads to an interesting topic: how to scale the system, especially when organizing such a large DRAM capacity, and why DDR4 was chosen over other options.<\/p>\n\n\n\n<p>According to the introduction, the SN10-8R has eight cards occupying 1\/4 of a rack. Each RDU contains six Channel Memory, with each channel having 256GB. An SN10-8R has 48 channels, totaling 12TB, using DDR4-2667 per channel, and the Q&amp;A also mentioned DDR4-3200.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/An-SN10-8R-has-48-channels.png\" alt=\"An SN10-8R has 48 channels\" class=\"wp-image-8287\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/An-SN10-8R-has-48-channels.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/An-SN10-8R-has-48-channels-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/An-SN10-8R-has-48-channels-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Furthermore, the software stack is compatible with mainstream AI frameworks. The Dataflow Graph Analyzer constructs sub-graphs suitable for scheduling and deployment on the RDU from the original DAG, forming the Dataflow Graph.<\/p>\n\n\n\n<p>The Dataflow Compiler compiles and deploys based on the Dataflow Graph, invoking the operator template library (Spatial Templates), similar to the P&amp;R tool in EDA, deploying the Dataflow Graph onto the RDU&#8217;s hardware computing and storage units.<\/p>\n\n\n\n<p>Additionally, SambaNova provides a way for users to customize operators, supporting user-defined operators based on their needs. The author speculates that this is based on the C+Intrinsic API provided by LLVM, as the SN10 core architecture includes many SIMD instructions, likely requiring manual coding.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambaflow-software.png\" alt=\"sambaflow software\" class=\"wp-image-8288\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambaflow-software.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambaflow-software-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambaflow-software-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>On reconfigurable hardware, communication is also programmable, a point strongly emphasized by another architecture, Groq.<\/p>\n\n\n\n<p>Compared to traditional architectures, the author&#8217;s personal understanding is that the benefits of this approach in SambaNova&#8217;s architecture are:<\/p>\n\n\n\n<p>1. The Spatial programming paradigm has more regular communication patterns, which can be software-controlled. In contrast, traditional approaches have more chaotic communication patterns, leading to potential congestion.<\/p>\n\n\n\n<p>2. With an enormous DRAM capacity, more data can be stored on-card, converting some inter-card communication to on-card DRAM data exchange, reducing communication overhead and lowering the demand for inter-card communication bandwidth.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambaflow-produces-highly-optimized-dataflow-mappings.png\" alt=\"sambaflow produces highly optimized dataflow mappings\" class=\"wp-image-8289\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambaflow-produces-highly-optimized-dataflow-mappings.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambaflow-produces-highly-optimized-dataflow-mappings-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/sambaflow-produces-highly-optimized-dataflow-mappings-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>GPUs excel at certain specific applications, but for very large models that cannot fit or run efficiently on a single GPU, their performance can suffer.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/goldilocks-zone.png\" alt=\"goldilocks zone\" class=\"wp-image-8290\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/goldilocks-zone.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/goldilocks-zone-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/goldilocks-zone-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>SambaNova&#8217;s dataflow architecture features numerous compute PCUs (0.5 TOPS) and storage PMUs (0.5M), which can be leveraged for parallel and data locality optimizations. The example below illustrates how chaining multiple PCUs enables fine-grained pipeline parallelism, with PMUs facilitating data exchange between PCUs and providing ideal data locality control.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/dataflow-exploits-data-locality-and-parallelism.png\" alt=\"dataflow exploits data locality and parallelism\" class=\"wp-image-8327\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/dataflow-exploits-data-locality-and-parallelism.png 832w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/dataflow-exploits-data-locality-and-parallelism-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/dataflow-exploits-data-locality-and-parallelism-768x432.png 768w\" sizes=\"(max-width: 832px) 100vw, 832px\" \/><\/figure>\n\n\n\n<p><strong>Potential Issues<\/strong><strong><\/strong><\/p>\n\n\n\n<p>Regular Data Parallelism: Operations like Parallel.For require a shared memory for data exchange. In SN10 and SN30, is this limited to DDR? In SN40L, HBM can be used, potentially improving latency.<\/p>\n\n\n\n<p>PCU and PMU Granularity: The granularity of PCUs and PMUs is very fine. How is the DAG decomposed to this level?<\/p>\n\n\n\n<p>(A) Does the Dataflow Graph Analyzer directly decompose to this fine granularity, with operators then interfacing with it?<\/p>\n\n\n\n<p>(B) Or does the Dataflow Graph Analyzer decompose to a coarser granularity, with operators controlling hardware resources and a coarse-grained operator being hand-written to interface?<\/p>\n\n\n\n<p>In my opinion, the graph layer provides a global view, enabling more optimization opportunities. However, a powerful automatic operator generation tool is needed to codegen operators from the subgraphs produced by the Dataflow Graph Analyzer using minimal resources. Hand-writing operators would be challenging for flexible graph-layer compilation. I speculate that the real-world approach might involve hand-writing operators for fixed patterns, with the graph layer&#8217;s Dataflow Graph Analyzer compiling based on these existing patterns, similar to most approaches.<\/p>\n\n\n\n<p>Fully exploiting the performance potential is a significant challenge. SambaNova&#8217;s architecture allows hardware reconfiguration to adapt to programs, unlike traditional architectures where programs must adapt to hardware. Combined with fine-grained computing and storage, this provides a larger tuning space.<\/p>\n\n\n\n<p>The SN10 architecture, as shown in the image, consists of Tiles made up of programmable interconnects, computing, and memory. Tiles can form larger logical Tiles, operate independently, enable data parallelism, or run different programs. Tiles access TB-level DDR4 memory, host machine memory, and support multi-card interconnects.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/cardinal-sn10.png\" alt=\"cardinal sn10\" class=\"wp-image-8328\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/cardinal-sn10.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/cardinal-sn10-300x160.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/cardinal-sn10-768x409.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Tiles in the SambaNova architecture are composed of three programmable units: Switch, PCU (Programmable Compute Unit), and PMU (Programmable Memory Unit). Understanding how to control these units is crucial for engineers writing operators.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Tiles-in-the-SambaNova.png\" alt=\"Tiles in the SambaNova\" class=\"wp-image-8329\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Tiles-in-the-SambaNova.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Tiles-in-the-SambaNova-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/Tiles-in-the-SambaNova-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>The PCU is primarily composed of configurable SIMD instructions. Programmable counters to execute inner loops in programs. A Tail Unit to accelerate transcendental functions like Sigmoid Funtion.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/The-PCU-is-primarily-composed-of-configurable-SIMD-instructions.png\" alt=\"The PCU is primarily composed of configurable SIMD instructions\" class=\"wp-image-8330\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/The-PCU-is-primarily-composed-of-configurable-SIMD-instructions.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/The-PCU-is-primarily-composed-of-configurable-SIMD-instructions-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/The-PCU-is-primarily-composed-of-configurable-SIMD-instructions-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>The PMU is a multi-banked SRAM that supports tensor data format conversions, such as&nbsp;Transpose.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/The-PMU-is-a-multi-banked-SRAM-that-supports-tensor-data-format-conversions-such-as-Transpose.png\" alt=\"The PMU is a multi-banked SRAM that supports tensor data format conversions, such as Transpose\" class=\"wp-image-8331\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/The-PMU-is-a-multi-banked-SRAM-that-supports-tensor-data-format-conversions-such-as-Transpose.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/The-PMU-is-a-multi-banked-SRAM-that-supports-tensor-data-format-conversions-such-as-Transpose-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/The-PMU-is-a-multi-banked-SRAM-that-supports-tensor-data-format-conversions-such-as-Transpose-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>The Switch contains a Router Pipeline and a Router Crossbar, controlling data flow between PCUs and PMUs. It includes a counter to implement outer loops.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/switch-and-on-chip-interconnect.png\" alt=\"switch and on-chip interconnect\" class=\"wp-image-8332\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/switch-and-on-chip-interconnect.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/switch-and-on-chip-interconnect-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/switch-and-on-chip-interconnect-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Here&#8217;s an example of using Layer Normalization, outlining the different computation steps. On the SambaNova hardware, the Switch connects different PCUs and PMUs to form pipelines. Nearby computations can rapidly exchange data through PMUs (a form of in-memory computing) while supporting concurrent pipelines (if the program has parallel DAGs).<\/p>\n\n\n\n<p>Note that this example follows a spatial programming paradigm, where different compute units execute unique computation programs (PCUs execute the same program at different times). Computations are fully unrolled in the spatial dimension, with some fusion applied.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/layernorm-pipelined-in-space.png\" alt=\"layernorm, pipelined in space\" class=\"wp-image-8333\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/layernorm-pipelined-in-space.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/layernorm-pipelined-in-space-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/layernorm-pipelined-in-space-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Building upon the Space, a temporal programming approach can be added to process data using minimal PMUs, MCUs, and Switches. The temporal approach considers executing specific programs on specific data at particular times. Intuitively, the same compute unit may execute different programs at different times. This approach uses fewer hardware resources, allowing the chip to execute more programs. Combined with high-speed on-chip interconnects and large DDR4 capacity, larger-scale programs can be processed. Both spatial and temporal programming modes are essential.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/hybrid-spacetime-execution.png\" alt=\"hybrid space+time execution\" class=\"wp-image-8334\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/hybrid-spacetime-execution.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/hybrid-spacetime-execution-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/hybrid-spacetime-execution-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>The following is a coarse-grained execution process for a small network, while the previous example focused on a Layer Normalization operator. Unlike other architectures (TPUs, Davinci), the SambaNova architecture natively supports on-chip fusion.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/spatial-dataflow-within-an-rdu.png\" alt=\"spatial dataflow within an rdu\" class=\"wp-image-8335\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/spatial-dataflow-within-an-rdu.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/spatial-dataflow-within-an-rdu-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/spatial-dataflow-within-an-rdu-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/of-a-rack-providing-12TB-of-DRAM-1.png\" alt=\"of a rack providing 12TB of DRAM\" class=\"wp-image-8338\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/of-a-rack-providing-12TB-of-DRAM-1.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/of-a-rack-providing-12TB-of-DRAM-1-300x152.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/of-a-rack-providing-12TB-of-DRAM-1-768x390.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>With 1\/4 of a rack providing 12TB of DRAM, there is sufficient capacity to fully utilize the chip&#8217;s compute power. The SN40L further adds 64GB of HBM, likely placed between the DDR4 and PMU SRAM based on the PR content.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/terabyte-sized-models.png\" alt=\"terabyte sized models\" class=\"wp-image-8339\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/terabyte-sized-models.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/terabyte-sized-models-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/terabyte-sized-models-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Such an architecture requires extremely powerful compilation techniques to support it, heavily relying on the compiler to map the dataflow graph onto the hardware resources. The following are various deployment scenarios that require compiler support.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/flexibility-to-support-key-scenarios.png\" alt=\"flexibility to support key scenarios\" class=\"wp-image-8340\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/flexibility-to-support-key-scenarios.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/flexibility-to-support-key-scenarios-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/flexibility-to-support-key-scenarios-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Multi-Card and Multi-Machine Interconnected Systems for Training and Inference<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/datascale-systems-scale-out.png\" alt=\"datascale systems scale-out\" class=\"wp-image-8341\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/datascale-systems-scale-out.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/datascale-systems-scale-out-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/datascale-systems-scale-out-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>The benefit of using a large DRAM capacity is that it can support training 1 Trillion Parameter Language Models without the need for model partitioning, unlike GPU systems. Here&#8217;s a comparison: training a 1T model on GPUs requires complex model parallelism techniques, while on the SambaNova system, the footprint is smaller. This raises an interesting point: Cerebras&#8217; Dojo wafer-scale technology also aims to reduce the compute footprint. Perhaps building upon SambaNova&#8217;s current work, their cores could be implemented using a wafer-scale approach similar to Cerebras Dojo (SambaNova&#8217;s PCUs and PMUs are also relatively fine-grained). Conversely, other architectures could consider SambaNova&#8217;s large DRAM capacity approach (these architectures also have large SRAMs and are dataflow-based).<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/train-large-language-models.png\" alt=\"train large language models\" class=\"wp-image-8342\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/train-large-language-models.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/train-large-language-models-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/train-large-language-models-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>The significance of reducing the footprint can be appreciated from Elon Musk&#8217;s statement. Enabling GPU clusters above 10K is a substantial undertaking, with difficulty and workload comparable to supercomputers, raising the technical barrier significantly. If the FSD 12 proves the viability of the end-to-end large model approach for autonomous driving, the availability of large compute clusters will become crucial for companies pursuing this technology.<\/p>\n\n\n\n<p>In the absence of H100 availability, meeting the training demands of large models for domestic leading automakers could quickly become a bottleneck.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/musk-x.png\" alt=\"musk x\" class=\"wp-image-8343\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/musk-x.png 598w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/musk-x-300x280.png 300w\" sizes=\"(max-width: 598px) 100vw, 598px\" \/><\/figure>\n\n\n\n<p>One of SambaNova&#8217;s other advantages is that it can support training on high-resolution images like 4K or even 50K x 50K without needing to downsample to a lower precision. If applied in the self-driving domain, training on high-definition video can provide accuracy advantages.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/4k-to-50k-convolutions.png\" alt=\"4k to 50k convolutions\" class=\"wp-image-8344\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/4k-to-50k-convolutions.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/4k-to-50k-convolutions-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/4k-to-50k-convolutions-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>On GPUs, downsampling is generally required (possibly first on the host CPU) to fit the data into the GPU&#8217;s DRAM. SambaNova&#8217;s large memory capacity means no downsampling or splitting is needed.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/enabling-high-resolution-full-image-pathology.png\" alt=\"enabling high-resolution full-image pathology\" class=\"wp-image-8345\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/enabling-high-resolution-full-image-pathology.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/enabling-high-resolution-full-image-pathology-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/enabling-high-resolution-full-image-pathology-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Supporting training on higher resolution images implies a need for higher training precision.<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/world-record-accuracy-high-res-convolution-training.png\" alt=\"world record accuracy high-res convolution training\" class=\"wp-image-8346\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/world-record-accuracy-high-res-convolution-training.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/world-record-accuracy-high-res-convolution-training-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/world-record-accuracy-high-res-convolution-training-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<p>Similarly, for recommendation models, accuracy improvements translate to massive economic gains. SambaNova&#8217;s large memory architecture is highly beneficial for boosting recommendation model accuracy.<\/p>\n\n\n\n<p>Here is SambaNova&#8217;s acceleration effect across different model scales:<\/p>\n\n\n\n<figure class=\"wp-block-image aligncenter size-full is-resized\"><img decoding=\"async\" src=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/surpassing-state-of-the-art-accuracy.png\" alt=\"surpassing state-of-the-art accuracy\" class=\"wp-image-8347\" style=\"width:800px\" width=\"800\" srcset=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/surpassing-state-of-the-art-accuracy.png 800w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/surpassing-state-of-the-art-accuracy-300x169.png 300w, https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/surpassing-state-of-the-art-accuracy-768x432.png 768w\" sizes=\"(max-width: 800px) 100vw, 800px\" \/><\/figure>\n\n\n\n<div id=\"ez-toc-container\" class=\"ez-toc-v2_0_76 ez-toc-wrap-left counter-hierarchy ez-toc-counter ez-toc-grey ez-toc-container-direction\">\n<div class=\"ez-toc-title-container\">\n<p class=\"ez-toc-title\" style=\"cursor:inherit\">Table of Contents<\/p>\n<span class=\"ez-toc-title-toggle\"><a href=\"#\" class=\"ez-toc-pull-right ez-toc-btn ez-toc-btn-xs ez-toc-btn-default ez-toc-toggle\" aria-label=\"Toggle Table of Content\"><span class=\"ez-toc-js-icon-con\"><span class=\"\"><span class=\"eztoc-hide\" style=\"display:none;\">Toggle<\/span><span class=\"ez-toc-icon-toggle-span\"><svg style=\"fill: #999;color:#999\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" class=\"list-377408\" width=\"20px\" height=\"20px\" viewBox=\"0 0 24 24\" fill=\"none\"><path d=\"M6 6H4v2h2V6zm14 0H8v2h12V6zM4 11h2v2H4v-2zm16 0H8v2h12v-2zM4 16h2v2H4v-2zm16 0H8v2h12v-2z\" fill=\"currentColor\"><\/path><\/svg><svg style=\"fill: #999;color:#999\" class=\"arrow-unsorted-368013\" xmlns=\"http:\/\/www.w3.org\/2000\/svg\" width=\"10px\" height=\"10px\" viewBox=\"0 0 24 24\" version=\"1.2\" baseProfile=\"tiny\"><path d=\"M18.2 9.3l-6.2-6.3-6.2 6.3c-.2.2-.3.4-.3.7s.1.5.3.7c.2.2.4.3.7.3h11c.3 0 .5-.1.7-.3.2-.2.3-.5.3-.7s-.1-.5-.3-.7zM5.8 14.7l6.2 6.3 6.2-6.3c.2-.2.3-.5.3-.7s-.1-.5-.3-.7c-.2-.2-.4-.3-.7-.3h-11c-.3 0-.5.1-.7.3-.2.2-.3.5-.3.7s.1.5.3.7z\"\/><\/svg><\/span><\/span><\/span><\/a><\/span><\/div>\n<nav><ul class='ez-toc-list ez-toc-list-level-1 ' ><li class='ez-toc-page-1 ez-toc-heading-level-2'><a class=\"ez-toc-link ez-toc-heading-1\" href=\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\/#Summary\" >Summary<\/a><ul class='ez-toc-list-level-3' ><li class='ez-toc-heading-level-3'><a class=\"ez-toc-link ez-toc-heading-2\" href=\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\/#Related_Posts\" >Related Posts<\/a><\/li><\/ul><\/li><\/ul><\/nav><\/div>\n<h2 class=\"wp-block-heading\" id=\"h-summary\"><span class=\"ez-toc-section\" id=\"Summary\"><\/span><strong>Summary<\/strong><strong><\/strong><span class=\"ez-toc-section-end\"><\/span><\/h2>\n\n\n\n<p>The RDU&#8217;s Dataflow architecture enables model execution on SRAM, reducing DRAM bandwidth demands. However, it also requires the compiler to deploy the Dataflow Graph onto the RDUs.<\/p>\n\n\n\n<p>The compiler uses a combined time+space technique to maximize hardware utilization for larger programs. Combined with high DRAM capacity, this supports larger models and high-resolution image training.<\/p>\n\n\n\n<p>Interpreting the SN40L Architecture Changes<\/p>\n\n\n\n<p>&#8220;SambaNova&#8217;s SN40L chip is unique. It addresses both HBM (High Bandwidth Memory) and DRAM from a single chip, enabling AI algorithms to choose the most appropriate memory for the task at hand, giving them direct access to far larger amounts of memory than can be achieved otherwise. Plus, by using SambaNova&#8217;s RDU (Reconfigurable Data Unit) architecture, the chips are designed to efficiently run sparse models using smarter compute.&#8221;<\/p>\n\n\n\n<p>The new SN40L architecture released in early September 2023 added more compute cores and introduced HBM for the first time. This section attempts to interpret the reasoning behind this choice.<\/p>\n\n\n\n<p>For inference, large language model parameters are tens to hundreds of GBs, and the KV cache is also tens to hundreds of GBs, while a single SN40 card&#8217;s SRAM is just a few hundred MBs (the SN10 was 320MB).<\/p>\n\n\n\n<p>For example, with the 65B parameter LLaMA model, storing the weights in FP16 format requires 65G*2=130GB. Taking LLaMA-65B as an example again, if the inference token length reaches the maximum of 2048 allowed by the model, the total KV cache size will be as high as 170GB.<\/p>\n\n\n\n<p>Ideally, parameters would reside in SRAM, the KV cache in memory close to compute, and inputs would stream directly on-chip to drive the compute pipeline. Obviously, a few hundred MB of SRAM per card is insufficient to host 100+GB of parameters while leaving room for inputs, even with multiple interconnected chips requiring thousands of chips. To reduce the footprint, a time-multiplexed programming model is still needed to reuse on-chip compute resources, loading different programs at different times, requiring weight streaming with different weights for different layers. However, if weights and the KV cache are stored in DDR, the latency would be quite high, making HBM an excellent medium for hosting hot data (weights, KV cache, or intermediate data that needs to spill).<\/p>\n\n\n\n<p>DDR is used for parameters and input data, while HBM acts as an ultra-large cache for hot data (weights, KV cache, or intermediate data about to be used).<\/p>\n\n\n\n<p>For fine-tuning and training, the bandwidth requirements are even greater. Introducing HBM provides increased bandwidth and capacity benefits.<\/p>\n<style>\r\n\r\n        .lwrp.link-whisper-related-posts{\r\n            \r\n            margin-top: 40px;\nmargin-bottom: 30px;\r\n        }\r\n        .lwrp .lwrp-title{\r\n            \r\n            \r\n        }\r\n        .lwrp .lwrp-description{\r\n            \r\n            \r\n\r\n        }\r\n        .lwrp .lwrp-list-container{\r\n        }\r\n        .lwrp .lwrp-list-multi-container{\r\n            display: flex;\r\n        }\r\n        .lwrp .lwrp-list-double{\r\n            width: 48%;\r\n        }\r\n        .lwrp .lwrp-list-triple{\r\n            width: 32%;\r\n        }\r\n        .lwrp .lwrp-list-row-container{\r\n            display: flex;\r\n            justify-content: space-between;\r\n        }\r\n        .lwrp .lwrp-list-row-container .lwrp-list-item{\r\n            width: calc(100% - 20px);\r\n        }\r\n        .lwrp .lwrp-list-item:not(.lwrp-no-posts-message-item){\r\n            \r\n            list-style: decimal;\r\n        }\r\n        .lwrp .lwrp-list-item img{\r\n            max-width: 100%;\r\n            height: auto;\r\n        }\r\n        .lwrp .lwrp-list-item.lwrp-empty-list-item{\r\n            background: initial !important;\r\n        }\r\n        .lwrp .lwrp-list-item .lwrp-list-link .lwrp-list-link-title-text,\r\n        .lwrp .lwrp-list-item .lwrp-list-no-posts-message{\r\n            \r\n                \r\n        }\r\n        @media screen and (max-width: 480px) {\r\n            .lwrp.link-whisper-related-posts{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-title{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-description{\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-list-multi-container{\r\n                flex-direction: column;\r\n            }\r\n            .lwrp .lwrp-list-multi-container ul.lwrp-list{\r\n                margin-top: 0px;\r\n                margin-bottom: 0px;\r\n                padding-top: 0px;\r\n                padding-bottom: 0px;\r\n            }\r\n            .lwrp .lwrp-list-double,\r\n            .lwrp .lwrp-list-triple{\r\n                width: 100%;\r\n            }\r\n            .lwrp .lwrp-list-row-container{\r\n                justify-content: initial;\r\n                flex-direction: column;\r\n            }\r\n            .lwrp .lwrp-list-row-container .lwrp-list-item{\r\n                width: 100%;\r\n            }\r\n            .lwrp .lwrp-list-item:not(.lwrp-no-posts-message-item){\r\n                \r\n                \r\n            }\r\n            .lwrp .lwrp-list-item .lwrp-list-link .lwrp-list-link-title-text,\r\n            .lwrp .lwrp-list-item .lwrp-list-no-posts-message{\r\n                \r\n                    \r\n            }\r\n        }<\/style>\r\n<div id=\"link-whisper-related-posts-widget\" class=\"link-whisper-related-posts lwrp\">\r\n            <h3 class=\"lwrp-title\">Related Posts<\/h3>    \r\n        <div class=\"lwrp-list-container\">\r\n                                            <ul class=\"lwrp-list lwrp-list-single\">\r\n                    <li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/gpu-clusters.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Unlocking the Potential of GPU Clusters for Advanced Machine Learning and Deep Learning Applications<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/questions\/when-100-swdm4-srbd-be-used.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Both the 100G-SWDM4 and the 100G-SRBD Transceivers Support 100G over Duplex Multi-Mode Fiber. When Should Each Transceiver Be Used?<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/400g-qsfp-dd-sr8-optical-transceiver.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">On 400G QSFP-DD SR8 Optical Transceiver Module<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/video\/how-to-use-400g-osfp-dr4-sr4-flat-top.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">How to Use 400G OSFP DR4 and SR4 Flat Top<\/span><\/a><\/li><li class=\"lwrp-list-item\"><a href=\"https:\/\/www.fibermall.com\/blog\/dgx-gh200.htm\" class=\"lwrp-list-link\"><span class=\"lwrp-list-link-title-text\">Revolutionizing AI: The NVIDIA DGX GH200 AI Supercomputer<\/span><\/a><\/li>                <\/ul>\r\n                        <\/div>\r\n<\/div>","protected":false},"excerpt":{"rendered":"<p>SambaNova Systems hailed as one of the top ten unicorn companies in the United States, raised $678 million in a Series D funding round led by SoftBank in April 2021, achieving a staggering $50 billion valuation. The company&#8217;s previous funding rounds involved prominent investors such as Google Ventures, Intel Capital, SK, Samsung Catalyst Fund, and [&hellip;]<\/p>\n","protected":false},"author":8,"featured_media":8286,"comment_status":"open","ping_status":"closed","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"site-sidebar-layout":"default","site-content-layout":"","ast-site-content-layout":"default","site-content-style":"default","site-sidebar-style":"default","ast-global-header-display":"","ast-banner-title-visibility":"","ast-main-header-display":"","ast-hfb-above-header-display":"","ast-hfb-below-header-display":"","ast-hfb-mobile-header-display":"","site-post-title":"","ast-breadcrumbs-content":"","ast-featured-img":"","footer-sml-layout":"","theme-transparent-header-meta":"","adv-header-id-meta":"","stick-header-meta":"","header-above-stick-meta":"","header-main-stick-meta":"","header-below-stick-meta":"","astra-migrate-meta-layouts":"set","ast-page-background-enabled":"default","ast-page-background-meta":{"desktop":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"ast-content-background-meta":{"desktop":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"tablet":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""},"mobile":{"background-color":"var(--ast-global-color-5)","background-image":"","background-repeat":"repeat","background-position":"center center","background-size":"auto","background-attachment":"scroll","background-type":"","background-media":"","overlay-type":"","overlay-color":"","overlay-opacity":"","overlay-gradient":""}},"footnotes":"","_wpscppro_dont_share_socialmedia":false,"_wpscppro_custom_social_share_image":0,"_facebook_share_type":"","_twitter_share_type":"","_linkedin_share_type":"","_pinterest_share_type":"","_linkedin_share_type_page":"","_instagram_share_type":"","_medium_share_type":"","_threads_share_type":"","_google_business_share_type":"","_selected_social_profile":[],"_wpsp_enable_custom_social_template":false,"_wpsp_social_scheduling":{"enabled":false,"datetime":null,"platforms":[],"status":"template_only","dateOption":"today","timeOption":"now","customDays":"","customHours":"","customDate":"","customTime":"","schedulingType":"absolute"},"_wpsp_active_default_template":true},"categories":[2,29],"tags":[],"class_list":["post-8282","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-blog","category-networking"],"acf":[],"yoast_head":"<!-- This site is optimized with the Yoast SEO Premium plugin v20.13 (Yoast SEO v25.8) - https:\/\/yoast.com\/wordpress\/plugins\/seo\/ -->\n<title>How does DSA Outperform NVIDIA GPUs? | FiberMall<\/title>\n<meta name=\"description\" content=\"According to SambaNova&#039;s early marketing materials, the company has taken a different approach to challenge the AI giant NVIDIA.\" \/>\n<meta name=\"robots\" content=\"index, follow, max-snippet:-1, max-image-preview:large, max-video-preview:-1\" \/>\n<link rel=\"canonical\" href=\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\" \/>\n<meta property=\"og:locale\" content=\"en_US\" \/>\n<meta property=\"og:type\" content=\"article\" \/>\n<meta property=\"og:title\" content=\"How does DSA Outperform NVIDIA GPUs?\" \/>\n<meta property=\"og:description\" content=\"SambaNova Systems hailed as one of the top ten unicorn companies in the United States, raised $678 million in a Series D funding round led by SoftBank in\" \/>\n<meta property=\"og:url\" content=\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\" \/>\n<meta property=\"og:site_name\" content=\"fibermall.com\" \/>\n<meta property=\"article:published_time\" content=\"2024-05-28T07:40:23+00:00\" \/>\n<meta property=\"article:modified_time\" content=\"2024-05-28T07:41:35+00:00\" \/>\n<meta property=\"og:image\" content=\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png\" \/>\n\t<meta property=\"og:image:width\" content=\"800\" \/>\n\t<meta property=\"og:image:height\" content=\"450\" \/>\n\t<meta property=\"og:image:type\" content=\"image\/png\" \/>\n<meta name=\"author\" content=\"Felisac\" \/>\n<meta name=\"twitter:card\" content=\"summary_large_image\" \/>\n<meta name=\"twitter:label1\" content=\"Written by\" \/>\n\t<meta name=\"twitter:data1\" content=\"Felisac\" \/>\n\t<meta name=\"twitter:label2\" content=\"Est. reading time\" \/>\n\t<meta name=\"twitter:data2\" content=\"16 minutes\" \/>\n<script type=\"application\/ld+json\" class=\"yoast-schema-graph\">{\"@context\":\"https:\/\/schema.org\",\"@graph\":[{\"@type\":\"Article\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#article\",\"isPartOf\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\"},\"author\":{\"name\":\"Felisac\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/6c504b906bd2221757ca517a21eb21b3\"},\"headline\":\"How does DSA Outperform NVIDIA GPUs?\",\"datePublished\":\"2024-05-28T07:40:23+00:00\",\"dateModified\":\"2024-05-28T07:41:35+00:00\",\"mainEntityOfPage\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\"},\"wordCount\":2275,\"commentCount\":0,\"publisher\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#primaryimage\"},\"thumbnailUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png\",\"articleSection\":[\"Blog\",\"Networking\"],\"inLanguage\":\"en-US\"},{\"@type\":\"WebPage\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\",\"url\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\",\"name\":\"How does DSA Outperform NVIDIA GPUs? | FiberMall\",\"isPartOf\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#website\"},\"primaryImageOfPage\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#primaryimage\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#primaryimage\"},\"thumbnailUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png\",\"datePublished\":\"2024-05-28T07:40:23+00:00\",\"dateModified\":\"2024-05-28T07:41:35+00:00\",\"description\":\"According to SambaNova's early marketing materials, the company has taken a different approach to challenge the AI giant NVIDIA.\",\"breadcrumb\":{\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#breadcrumb\"},\"inLanguage\":\"en-US\",\"potentialAction\":[{\"@type\":\"ReadAction\",\"target\":[\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm\"]}]},{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#primaryimage\",\"url\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png\",\"contentUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png\",\"width\":800,\"height\":450,\"caption\":\"SambaNova SN10 RDU at Hot Chips 33\"},{\"@type\":\"BreadcrumbList\",\"@id\":\"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#breadcrumb\",\"itemListElement\":[{\"@type\":\"ListItem\",\"position\":1,\"name\":\"Home\",\"item\":\"https:\/\/www.fibermall.com\/blog.htm\"},{\"@type\":\"ListItem\",\"position\":2,\"name\":\"How does DSA Outperform NVIDIA GPUs?\"}]},{\"@type\":\"WebSite\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#website\",\"url\":\"https:\/\/www.fibermall.com\/blog.htm\/\",\"name\":\"fibermall.com\",\"description\":\"Optical Communication Expert\",\"publisher\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\"},\"potentialAction\":[{\"@type\":\"SearchAction\",\"target\":{\"@type\":\"EntryPoint\",\"urlTemplate\":\"https:\/\/www.fibermall.com\/blog.htm\/?s={search_term_string}\"},\"query-input\":{\"@type\":\"PropertyValueSpecification\",\"valueRequired\":true,\"valueName\":\"search_term_string\"}}],\"inLanguage\":\"en-US\"},{\"@type\":\"Organization\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#organization\",\"name\":\"fibermall.com\",\"url\":\"https:\/\/www.fibermall.com\/blog.htm\/\",\"logo\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/\",\"url\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg\",\"contentUrl\":\"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg\",\"width\":200,\"height\":67,\"caption\":\"fibermall.com\"},\"image\":{\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/\"}},{\"@type\":\"Person\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/6c504b906bd2221757ca517a21eb21b3\",\"name\":\"Felisac\",\"image\":{\"@type\":\"ImageObject\",\"inLanguage\":\"en-US\",\"@id\":\"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/image\/\",\"url\":\"https:\/\/secure.gravatar.com\/avatar\/0fee6bec0831365a6e53fa374f659c986987cc7c085842868236d48e2d5f048c?s=96&d=mm&r=g\",\"contentUrl\":\"https:\/\/secure.gravatar.com\/avatar\/0fee6bec0831365a6e53fa374f659c986987cc7c085842868236d48e2d5f048c?s=96&d=mm&r=g\",\"caption\":\"Felisac\"},\"description\":\"Optical Technician\",\"sameAs\":[\"https:\/\/www.fibermall.com\/\"],\"url\":\"https:\/\/www.fibermall.com\/blog.htm?author=8\"}]}<\/script>\n<!-- \/ Yoast SEO Premium plugin. -->","yoast_head_json":{"title":"How does DSA Outperform NVIDIA GPUs? | FiberMall","description":"According to SambaNova's early marketing materials, the company has taken a different approach to challenge the AI giant NVIDIA.","robots":{"index":"index","follow":"follow","max-snippet":"max-snippet:-1","max-image-preview":"max-image-preview:large","max-video-preview":"max-video-preview:-1"},"canonical":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm","og_locale":"en_US","og_type":"article","og_title":"How does DSA Outperform NVIDIA GPUs?","og_description":"SambaNova Systems hailed as one of the top ten unicorn companies in the United States, raised $678 million in a Series D funding round led by SoftBank in","og_url":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm","og_site_name":"fibermall.com","article_published_time":"2024-05-28T07:40:23+00:00","article_modified_time":"2024-05-28T07:41:35+00:00","og_image":[{"width":800,"height":450,"url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png","type":"image\/png"}],"author":"Felisac","twitter_card":"summary_large_image","twitter_misc":{"Written by":"Felisac","Est. reading time":"16 minutes"},"schema":{"@context":"https:\/\/schema.org","@graph":[{"@type":"Article","@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#article","isPartOf":{"@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm"},"author":{"name":"Felisac","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/6c504b906bd2221757ca517a21eb21b3"},"headline":"How does DSA Outperform NVIDIA GPUs?","datePublished":"2024-05-28T07:40:23+00:00","dateModified":"2024-05-28T07:41:35+00:00","mainEntityOfPage":{"@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm"},"wordCount":2275,"commentCount":0,"publisher":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization"},"image":{"@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#primaryimage"},"thumbnailUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png","articleSection":["Blog","Networking"],"inLanguage":"en-US"},{"@type":"WebPage","@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm","url":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm","name":"How does DSA Outperform NVIDIA GPUs? | FiberMall","isPartOf":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#website"},"primaryImageOfPage":{"@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#primaryimage"},"image":{"@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#primaryimage"},"thumbnailUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png","datePublished":"2024-05-28T07:40:23+00:00","dateModified":"2024-05-28T07:41:35+00:00","description":"According to SambaNova's early marketing materials, the company has taken a different approach to challenge the AI giant NVIDIA.","breadcrumb":{"@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#breadcrumb"},"inLanguage":"en-US","potentialAction":[{"@type":"ReadAction","target":["https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm"]}]},{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#primaryimage","url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png","contentUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/05\/SambaNova-SN10-RDU-at-Hot-Chips-33.png","width":800,"height":450,"caption":"SambaNova SN10 RDU at Hot Chips 33"},{"@type":"BreadcrumbList","@id":"https:\/\/www.fibermall.com\/blog\/how-does-dsa-outperform-nvidia-gpus.htm#breadcrumb","itemListElement":[{"@type":"ListItem","position":1,"name":"Home","item":"https:\/\/www.fibermall.com\/blog.htm"},{"@type":"ListItem","position":2,"name":"How does DSA Outperform NVIDIA GPUs?"}]},{"@type":"WebSite","@id":"https:\/\/www.fibermall.com\/blog.htm\/#website","url":"https:\/\/www.fibermall.com\/blog.htm\/","name":"fibermall.com","description":"Optical Communication Expert","publisher":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization"},"potentialAction":[{"@type":"SearchAction","target":{"@type":"EntryPoint","urlTemplate":"https:\/\/www.fibermall.com\/blog.htm\/?s={search_term_string}"},"query-input":{"@type":"PropertyValueSpecification","valueRequired":true,"valueName":"search_term_string"}}],"inLanguage":"en-US"},{"@type":"Organization","@id":"https:\/\/www.fibermall.com\/blog.htm\/#organization","name":"fibermall.com","url":"https:\/\/www.fibermall.com\/blog.htm\/","logo":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/","url":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg","contentUrl":"https:\/\/www.fibermall.com\/blog\/wp-content\/uploads\/2024\/12\/cropped-Fiber-Mall-Logo.jpg","width":200,"height":67,"caption":"fibermall.com"},"image":{"@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/logo\/image\/"}},{"@type":"Person","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/6c504b906bd2221757ca517a21eb21b3","name":"Felisac","image":{"@type":"ImageObject","inLanguage":"en-US","@id":"https:\/\/www.fibermall.com\/blog.htm\/#\/schema\/person\/image\/","url":"https:\/\/secure.gravatar.com\/avatar\/0fee6bec0831365a6e53fa374f659c986987cc7c085842868236d48e2d5f048c?s=96&d=mm&r=g","contentUrl":"https:\/\/secure.gravatar.com\/avatar\/0fee6bec0831365a6e53fa374f659c986987cc7c085842868236d48e2d5f048c?s=96&d=mm&r=g","caption":"Felisac"},"description":"Optical Technician","sameAs":["https:\/\/www.fibermall.com\/"],"url":"https:\/\/www.fibermall.com\/blog.htm?author=8"}]}},"_links":{"self":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/8282","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/users\/8"}],"replies":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fcomments&post=8282"}],"version-history":[{"count":6,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/8282\/revisions"}],"predecessor-version":[{"id":8363,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/posts\/8282\/revisions\/8363"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=\/wp\/v2\/media\/8286"}],"wp:attachment":[{"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fmedia&parent=8282"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Fcategories&post=8282"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/www.fibermall.com\/blog.htm\/index.php?rest_route=%2Fwp%2Fv2%2Ftags&post=8282"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}