CV
Professional CV covering robotics, computer vision, research, engineering experience, and technical projects.
Contact Information
| Name | Yue Xin (辛约) |
| Professional Title | Ph.D. Student | Robot Learning, Embodied AI, and Intelligent Systems |
| xinyue4496@163.com | |
| Location | Beijing / Shanghai |
| Website | https://iconssss.github.io |
Professional Summary
Direct Ph.D. student at Tsinghua University with interdisciplinary training in physics, mathematics, measurement and control, and complex sensing systems. I combine rigorous modeling and experimental engineering with hands-on work in visuomotor policy learning, VLA training and evaluation, computer vision, 3D perception, ROS2 policy execution, and latency-aware robot systems. I am seeking robot learning and embodied AI research engineering roles that value both scientific discipline and end-to-end implementation.
Experience
-
2026.06 - 2026.07 China
Industry Practice — Computer Vision Algorithms and Edge Deployment
Goertek Inc.
Developed and delivered a multi-person motion perception and identity recognition system for intelligent-device interaction, covering the full workflow from requirements and data collection to evaluation, GPU optimization, Unity integration, and Windows edge deployment.
- Built a multi-person pose and motion-analysis pipeline with YOLO Pose and OpenCV, including ROI/slot-based perception and a causal temporal state machine for periodic-motion recognition.
- Implemented temporal keypoint fusion, scale normalization, motion gating, and event-level manual annotation and matching for systematic evaluation.
- Applied FP16 inference, temporal micro-batching, asynchronous video acquisition, and geometry-consistent preprocessing to achieve approximately 30 FPS video processing on an RTX 3050.
- Evaluated a three-person historical test video with 145 ground-truth events, 142 predictions, and 141 matches, corresponding to an approximately 97.2% event match rate.
- Built a local identity module with YuNet and SFace and established the binding workflow between student identity and visual slots.
- Integrated heartbeat, identity, task-control, and real-time result messages with a Unity client through UDP/OSC, then verified delivery on a Windows RTX 3050 edge workstation.
Education
-
2024 - Present Beijing, China
Ph.D. Student (Direct Ph.D. Track)
Department of Precision Instrument, Tsinghua University
Precision Instrument
- Researching complex inertial sensing systems through dynamic modeling, numerical simulation, error analysis, signal processing, parameter calibration, and experimental platform integration.
- Developed experience in reasoning about sensing uncertainty, system identification and debugging, and hardware-software interaction.
-
2020 - 2024 Beijing, China
Bachelor's Degree, Dual-degree Program
Weiyang College, Tsinghua University
Mathematical and Physical Basic Sciences; Measurement, Control Technology and Instruments
- Completed interdisciplinary training across physics, mathematics, numerical analysis, control, measurement, and instrumentation.
- Built a rigorous foundation for cross-domain research and engineering problem solving.
Skills
Languages
Projects
-
Reliable Multi-GPU VLA Fine-Tuning and Evaluation
Built an end-to-end SmolVLA training and evaluation workflow on LIBERO, from native observation/action contract validation to controlled stochastic-policy checkpoint comparison.
- Trained on 1,693 episodes / 273,465 frames with 4×RTX 4090 DDP, reaching 3.82× speedup and 289 samples/s at global batch 64.
- Integrated official LeRobot training and LIBERO closed-loop evaluation, including self-contained checkpoints and optimizer-state resume from 5K to 20K steps.
- Identified task-order-dependent flow-matching RNG as a checkpoint-comparison confound and implemented per-episode paired policy-noise control; the corrected 5K→20K result was 2/30→4/30, reported as a weak signal rather than a significance claim.
-
Latency-Aware Runtime for Closed-Loop Robot Policies
Designed an asynchronous learned-policy runtime that tracks action provenance and controls scheduler-induced staleness through explicit synchronous, FIFO, and latest-only semantics.
- Decomposed action age into observation wait, inference latency, and post-inference wait, with per-action timestamp and queue-depth instrumentation.
- On MuJoCo Reacher at 100 ms policy latency, LATEST reduced p95 action age from 1.24 s to 150 ms and improved success from 20% to 100% while preserving approximately 20 Hz control.
- Validated the mechanism on both a learned action-chunk benchmark and an independent MuJoCo environment with a fixed learned SAC policy.
-
Metric 3D Representations for Low-Data Visuomotor Control
Built a calibrated RGB-D to world-frame point-cloud pipeline and a compact PointNet policy, then mapped sample efficiency, viewpoint robustness, occlusion recovery, and calibration failure boundaries.
- Achieved 95% closed-loop success with 50 demonstrations and 100% with 250 demonstrations on a contact-free MuJoCo Panda reach task.
- Retained 96.7–100% success across 0–45° camera motion by reconstructing observations into a common world frame.
- Restored strong-occlusion performance from 3.3% to 100% through two-view fusion without policy retraining; quantified collapse under depth and extrinsic error.
-
XR-1 VLA Action-Prefix Temporal Alignment
Audited whether the native action-prefix interface of the 5.1B-parameter Xiaomi Robotics XR-1 VLA provides robustness to delayed asynchronous replanning.
- Implemented causal self-generated prefix plumbing for 30×60 action chunks and ran 190 formal generations across 4×RTX 4090.
- Measured 288 ms mean / 318 ms p95 generation latency and found growing right-arm handoff degradation at medium and large phase offsets.
- Demonstrated selective post-training feasibility for 675M parameters while reporting the study as a real-model systems boundary rather than a physical-robot success claim.
Research and Engineering Focus
- Visuomotor policy learning for manipulation, including imitation learning, action chunking, replanning, generalization, and closed-loop evaluation.
- VLA training and evaluation systems, with emphasis on model/data/action contracts, cross-embodiment transfer, checkpoint comparison, and failure diagnosis.
- Metric 3D and multi-view perception for robot interaction, including calibration sensitivity and robustness under distribution shift.
- Latency-aware robot policy runtimes, asynchronous inference semantics, ROS2 execution, motion planning, and reliable control.
- Reproducible embodied AI experimentation through controlled baselines, ablations, multi-seed evaluation, provenance tracking, and honest negative-result analysis.
Currently Exploring
- VLA and robot foundation model adaptation for manipulation and cross-embodiment transfer.
- Recovery-aware imitation learning, multimodal behavior, and flow-based robot policies.
- OpenPI and π0-family training, evaluation, and deployment workflows.
Professional Strengths
- Rigorous foundation in physics, mathematics, control, and system modeling, supported by Tsinghua dual-degree training.
- Ability to move between model formulation, software implementation, evaluation, and experimental or edge-system debugging.
- Hands-on experience across research prototypes, simulation benchmarks, GPU training systems, and delivered computer-vision applications.
- Transferable systems perspective connecting sensing uncertainty, control, policy learning, and embodied execution.
Honors and Awards
-
National High School Mathematics League (Shanghai Division), Second Prize
Competition Award
-
National High School Physics Competition (Shanghai Division), Second Prize
Competition Award