ZAO 2022 Zhongguancun Online Annual Observation Select Leading Solutions 30 (hereinafter referred to as LS30), providing industry users with better choices and helping industry high-quality solutions and technical solutions.
Zhongguancun Online believes that the Habana Gaudi 2 processor launched by Intel in 2022 can participate in this ultimate selection. The Habana Gaudi 2 processor adopts a 7-nanometer process technology, based on Habana's energy-efficient architecture, and is aimed at providing higher performance for customers' model training and inference.
●'s significance to data centers: helping to achieve efficient AI training
Nowadays, artificial intelligence is moving from a technical concept to a variety of industries and realizing the actual implementation of multiple scenarios. It can be seen that the artificial intelligence boom is promoting the continuous renewal of the AI chip market. According to Allied Market Research, the global machine learning chip market size will reach about US$37.8 billion by 2025. This not only drives the transformation of traditional chip companies' strategies and technologies, but also drives the entry of a large number of new players, and frequently makes surprise moves in continuity or disruptive innovation.
This year, Intel launched the Gaudi processor for high-performance deep learning AI training, which allows customers to train more at a lower cost. The newly released Habana Gaudi2 is developed based on the Synapse AI software stack. It can enable end users to make full use of the high performance and energy efficiency of the processor by supporting a diverse architecture.
For data centers, due to the increasing scale and complexity of data sets and artificial intelligence businesses, the time and cost of training deep learning models is getting higher and higher. According to IDCh data, among the machine learning practitioners surveyed in 2020, 74% of their models have undergone 5-10 iteration training, more than 50% need to reconstruct the models weekly or more frequently, and 26% will reconstruct the models every day or even hourly. 56% of respondents believe that the cost of training is the primary factor that hinders their organizations from using artificial intelligence to solve problems, innovate and enhance the end-customer experience. Intel's Habana Gaudi 2 processor adopts a 7nm process and is based on Habana's energy-efficient architecture, aimed at data center-oriented computer vision and natural language applications, aiming to provide higher performance for customers' model training and inference.
● technical analysis: All-round upgrades effectively improve training performance
Based on the same architecture as the first generation of Gaudi, the Habana Gaudi 2 processor greatly improves training performance. When customers run Amazon EC2 DL1 instances in the cloud and run Supermicro locally, their cost-effectiveness is 40% higher than the existing GPU solution. These are all due to the advancement of Gaudi2 in architecture: , including process technology, jumped from 16 nm to 7 nm; new data types including FP8 were introduced into the matrix multiplication (MME) and the Tensor processor core computing engine; the number of Tensor processor core cores has increased from 8 to 24; the multimedia processing engine on-chip is integrated to realize uninstallation from the host subsystem; the memory capacity of the on-chip package has increased by 3 times, from 32GB to 96GB with a bandwidth of 2.45TB/sec HBM2E; double 48MB of onboard SRAM memory and integrated Ethernet based on RDMA (RoCE2) increase from 10 to 24, enabling efficient vertical and horizontal scaling on standard networks.
can also be seen from the performance of MLPerf industry tests that the Habana Gaudi 2 processor has a considerable advantage in training time compared to the NVIDIA A100 on visual (ResNet-50) and language (BERT) models.
Compared with the first generation of Gaudi processors, the training throughput of Habana Gaudi 2 processor in the ResNet-50 model has been increased by 3 times, and the training throughput of BERT model has been increased by 4.7 times. These are attributed to the process process improvement from 16nm to 7nm, the number of Tensor processor cores tripled, the increase in GEMM engine computing power, the package's high-bandwidth storage capacity tripled, the SRAM bandwidth increased, and the doubled capacity. For the training of visual processing models, the Gaudi2 processor integrates a media processing engine, which can independently complete the preprocessing of data augmentation and compressed images required for AI training.
The performance of the two generations of Gaudi processors is achieved through the commercial software stack for Habana customers out of the box without special software operation.
is tested and compared on Habana 8 GPU servers and HLS-Gaudi2 reference servers through out-of-the-box performance provided by commercial software. Among them, the training throughput comes from the TensorFlow docker of NGC and Habana public libraries, and uses the best performance parameters recommended by both parties to measure in the mixed precision training mode. It is worth noting that throughput is a key factor affecting the convergence of the final training time.
●Industry impact and user needs: Data centers accelerate on demand, making deep learning "faster"
By deploying Habana Gaudi 2 to the data center, it can provide higher efficiency for model training and reasoning for computer vision and natural language processing, and solve two of the most concerned issues of customers: reducing server processing costs and reducing the time required to train models. Habana Gaudi2 and Greco AI accelerators are developed based on the Synapse AI software stack. By supporting a diverse architecture, end users can make full use of the high performance and energy efficiency of the processor.
At the same time, with the help of Habana Labs' Gaudi platform, the data center team can focus on deep learning processor technology, allowing data scientists and machine learning engineers to efficiently train models, and implement new model building or existing model migration through simple code, improving work efficiency while reducing operating costs .
●Conclusion
In response to the "basic computing power" field that mainly provides computing power for cloud computing , edge computing and other needs, Intel's second-generation Gaudi processor Habana Gaudi2 has achieved a key leap in deep learning. By supporting diversified architectures, users can make full use of the high performance and energy efficiency of the processor to train data center loads with higher cost performance. There is no doubt that in scenarios where server or server cluster is mainly used for deep learning training and inference computing, Habana Gaudi2 is an ideal accelerator for these dedicated scenarios, it can provide excellent deep learning performance and reduce the total cost of ownership.
(8086572)