WoRV
EngineeringFull-time

Inference Optimization Engineer

제2판교 IT센터
상시 채용
Apply전문연구요원 / 산업기능요원 지원 가능

About the Team

We enable Physical AI to operate reliably in the real world. Our mission is to maximize inference efficiency in on-device environments while meeting accuracy and latency requirements under real-world constraints. By understanding the computational structure of VLAs, building dedicated runtime engines, and engineering the software stack from scratch, we shape the next-generation of serving systems.

Responsibilities

  • Analyze the architecture and operations of generative models, especially VLAs
  • Deploy and run models on edge devices, utilizing hardware accelerators such as NPUs
  • Identify performance bottlenecks, then research and apply techniques to resolve them

Minimum Qualifications

  • Proficiency in C/C++ or Python
  • Fundamental CS knowledge and English proficiency to understand ML systems research papers
  • A proactive approach to problem-solving

Preferred Qualifications

  • Experience deploying and optimizing generative models on edge devices
  • Experience profiling performance at kernel and hardware level
  • Kernel programming experience on hardware accelerators such as GPUs or NPUs

Hiring Process

1
Application
2
1st Interview
Technical interview
3
Offer
4
Hired

We're looking for low-ego, passionate teammates to join us