Thank you for sending your enquiry! One of our team members will contact you shortly.
Thank you for sending your booking! One of our team members will contact you shortly.
Course Outline
The Landscape of Chinese AI GPUs
- Comparative analysis of Huawei Ascend, Biren, and Cambricon MLU
- Differences between CUDA and CANN, Biren SDK, and BANGPy paradigms
- Market trends and vendor ecosystems
Migration Preparation
- Conducting an assessment of your current CUDA codebase
- Determining target platforms and corresponding SDK versions
- Configuring toolchains and development environments
Techniques for Code Translation
- Adapting CUDA memory access patterns and kernel logic
- Translating compute grid and thread models
- Evaluating automated versus manual translation methods
Implementations Specific to Platforms
- Leveraging Huawei CANN operators and custom kernels
- Utilizing the Biren SDK conversion workflow
- Reconstructing models using BANGPy (Cambricon)
Testing and Optimization Across Platforms
- Profiling execution efficiency on each intended platform
- Optimizing memory usage and comparing parallel execution models
- Monitoring performance metrics and iterative refinement
Handling Heterogeneous GPU Environments
- Deploying hybrid systems with multiple architectural types
- Implementing fallback mechanisms and device identification
- Establishing abstraction layers to ensure code maintainability
Real-World Examples and Best Practices
- Porting vision or NLP models to Ascend or Cambricon hardware
- Adapting inference pipelines for Biren clusters
- Resolving version incompatibilities and API discrepancies
Conclusion and Forward-Looking Steps
Requirements
- Hands-on experience in programming with CUDA or GPU-accelerated applications
- Solid comprehension of GPU memory models and compute kernels
- Proficiency in AI model deployment or acceleration workflows
Target Audience
- GPU Developers
- System Architects
- Software Porting Specialists
21 Hours