Required Skills: computer vision algorithm design, spatial geometry, low-level C++, CUDA, Orin NX, AGX Orin, L4T, power profiles, core pinning, PTQ, QAT, TensorRT FP16/INT8 engines, Python, PyTorch, OpenCV, UDP, PLC integration
Job Description
Core Objective: Architect, benchmark, and optimize an end-to-end computer vision and low-latency execution pipeline—combining spatial geometric math with low-level C++/CUDA acceleration directly on NVIDIA Jetson embedded edge hardware (Orin NX / AGX Orin) to achieve sub-10 ms processing latency.
Key Responsibilities
1. Model Selection & Spatial Math: Evaluate real-time detection topologies (e.g. YOLO) at 100+ FPS and build 2D perspective homography unwarping, lens undistortion, and spatial algorithms to convert pixel coordinates into physical millimeter units within tight error bounds (millimeter level accuracy).
2. TensorRT INT8 Acceleration: Execute Post-Training Quantization to compile PyTorch/ONNX models into high-throughput TensorRT INT8/FP16 engines on Jetson Orin hardware without accuracy degradation.
3. Zero-Copy Memory Architecture: Engineer zero-copy memory pipelines using NVIDIA Memory Management and DMA transfers to eliminate bottlenecks.
4. Compiled C++ Execution & I/O: Build compiled C++17 execution frameworks with multi-threaded lock-free ring buffers and non-blocking asynchronous socket communication.
5. Nsight Profiling & Roadmapping: Instrument stage-by-stage pipeline latency using NVIDIA Nsight Systems/NVTX markers to bound tail latency, author feasibility reports, and design technical roadmaps.
Key Qualifications & Experience
1. Full-Stack Edge AI Experience: 12–15+ years of experience bridging real-time computer vision algorithm design, spatial geometry, and low-level C++/CUDA execution on embedded edge hardware.
2. NVIDIA Jetson Ecosystem: Deep expertise in Jetson embedded platforms (Orin NX, AGX Orin, L4T, power profiles, core pinning) and a proven track record compiling/tuning TensorRT FP16/INT8 engines via PTQ/QAT.
3. Optical & Spatial Geometry: Expertise in 2D/3D camera calibration, perspective homography transformations, lens distortion modeling, and millimeter sizing math in industrial settings.
4. Low-Latency Systems Engineering: Expertise in zero-copy shared memory, DMA frame buffers, lock-free queues, custom CUDA plugins, and microsecond profiling via Nsight Systems and NVTX markers.
5. Tooling & Industrial I/O: Proficiency in C++17, Python, PyTorch, OpenCV, CUDA, non-blocking asynchronous sockets (UDP, PLC integration), and building automated dataset benchmarking harnesses.