Get in Touch

Course Outline

The Landscape of Chinese AI GPUs

  • Comparative analysis of Huawei Ascend, Biren, and Cambricon MLU
  • Differences between CUDA and CANN, Biren SDK, and BANGPy paradigms
  • Market trends and vendor ecosystems

Migration Preparation

  • Conducting an assessment of your current CUDA codebase
  • Determining target platforms and corresponding SDK versions
  • Configuring toolchains and development environments

Techniques for Code Translation

  • Adapting CUDA memory access patterns and kernel logic
  • Translating compute grid and thread models
  • Evaluating automated versus manual translation methods

Implementations Specific to Platforms

  • Leveraging Huawei CANN operators and custom kernels
  • Utilizing the Biren SDK conversion workflow
  • Reconstructing models using BANGPy (Cambricon)

Testing and Optimization Across Platforms

  • Profiling execution efficiency on each intended platform
  • Optimizing memory usage and comparing parallel execution models
  • Monitoring performance metrics and iterative refinement

Handling Heterogeneous GPU Environments

  • Deploying hybrid systems with multiple architectural types
  • Implementing fallback mechanisms and device identification
  • Establishing abstraction layers to ensure code maintainability

Real-World Examples and Best Practices

  • Porting vision or NLP models to Ascend or Cambricon hardware
  • Adapting inference pipelines for Biren clusters
  • Resolving version incompatibilities and API discrepancies

Conclusion and Forward-Looking Steps

Requirements

  • Hands-on experience in programming with CUDA or GPU-accelerated applications
  • Solid comprehension of GPU memory models and compute kernels
  • Proficiency in AI model deployment or acceleration workflows

Target Audience

  • GPU Developers
  • System Architects
  • Software Porting Specialists
 21 Hours

Number of participants


Price per participant

Upcoming Courses

Related Categories