
GPU acceleration engineer
GECI Int. ยท Remote
About The Role
GPU Acceleration Engineer - Calculation Engine
๐ฏ Main Mission
Massively accelerate the sparse calculation engine of a UK SaaS B2B - Enterprise Planning & Analytics company by porting critical algorithms from Rust/C++ to GPU (CUDA). Transform currently impossible calculations (requiring thousands of years of CPU time) into operations achievable in minutes.
๐ Context
UK SaaS B2B - Enterprise Planning & Analytics company manages planning models reaching 64 quadrillion cells with billions of time periods. Our Hyperblock/Polaris engine is currently limited by:
- Legacy CPU architecture (Java/Rust/C++)
- Memory constraints on massive sparse structures
- Prohibitive calculation times on complex scenarios
- Objective : Achieve performance gains of 100x to 1000x via GPU offloading.
- ๐ง Main Responsibilities
GPU Offloading
- Port existing Rust/C++ algorithms to CUDA/GPU
- Identify and extract critical calculation paths to accelerate
- Optimize sparse matrix operations for GPU architecture
- Develop performant Rust โ CUDA wrappers
- Benchmark and validate performance gains
Memory Optimization
- Design GPU memory management strategies for massive datasets
- Implement efficient patterns for sparse structures
- Optimize CPU โ GPU memory transfers
- Manage GPU memory limitations on large-scale calculations
Technical Collaboration
- Work with engineering team on integration
- Document GPU porting patterns
- Participate in code reviews and design reviews
- Train the team on GPU best practices
- ๐ป Technical Stack
- Languages (in order of importance)
- CUDA - Primary GPU development
- Rust - Source language for algorithms to port
- C++ - Legacy components and CUDA interoperability
- (Java - platform context, no dev required)
Key Technologies
- NVIDIA CUDA (toolkit, libraries: cuBLAS, cuSPARSE)
- Rust (ownership model, unsafe blocks, FFI)
- GPU Programming (kernels, memory hierarchy, optimization)
- Sparse Matrix Operations (compression, storage formats)
- Profiling Tools (nvprof, Nsight, perf)
โ Required Profile
Essential Skills
GPU & CUDA (Essential)
- โ Significant CUDA programming experience (3+ years)
- โ Mastery of GPU kernel optimization
- โ Deep knowledge of NVIDIA GPU architecture (memory hierarchy, warps, occupancy)
- โ Experience with sparse calculations on GPU (cuSPARSE or equivalent)
Rust (Essential)
- โ Production Rust development
- โ Mastery of ownership and borrowing system
- โ Experience with unsafe Rust and FFI (Foreign Function Interface)
- โ Ability to analyze and refactor existing Rust code
C++ (Required)
- โ Modern C++ (C++11/14/17)
- โ C++ โ CUDA integration
- โ Templates and metaprogramming (asset)
Algorithms (Required)
- โ Data structures for scientific computing
- โ Sparse matrix algorithms (CSR, COO, etc.)
- โ Performance optimization and profiling
- โ Parallelization and concurrency concepts
Highly Valued Experience
- ๐ฏ Documented CPU โ GPU porting projects
- ๐ฏ HPC experience (supercomputers, GPU clusters)
- ๐ฏ Memory optimization for large-scale datasets
- ๐ฏ Scientific computing or numerical simulation
- ๐ฏ Rust interop with other languages (C/C++/Python)
๐ Working Arrangements
Location & Travel
- 100% remote (France/Europe base preferred)
- Occasional travel to London
- Frequency: ~1 week/month for team sprints
- Project kickoff + key reviews
- Intensive collaboration sessions
Start date : As soon as possible
This is an external listing. JobSpring does not represent or verify the employer. Report this listing
JobSpring