An End-to-End Stereo Vision Platform Enables Calibrated Depth and Multi-View 3D Reconstruction with Consumer Cameras
DOI:
https://doi.org/10.13021/jssr2026.5610Abstract
Stereo vision estimates depth from pixel displacement between paired images and supports applications including robotics, digital twins, and industrial inspection. Although modern stereo models produce dense disparity maps, deploying them with physical cameras still requires calibration, continuous capture, metric-depth conversion, quality control, and multi-view fusion. We developed a unified stereo-vision platform that integrates pretrained FoundationStereo, CREStereo, and IGEV++ models with consumer camera hardware through guided calibration, persistent GPU inference, confidence filtering, point-cloud generation, and video-to-3D reconstruction. Calibration using an 8×5 internal-corner checkerboard with 25 mm square pitch achieved a 0.427-pixel stereo reprojection error and recovered a 60.4 mm baseline. A browser interface manages camera capture, calibration, benchmarking, local or cluster execution, and artifact export from stereo image pairs, recordings, and multi-view captures. FoundationStereo evaluation produced 1.329% BP-2 on all 15 Middlebury v3 quarter-resolution pairs and 0.487% BP-1 on all 27 ETH3D pairs, closely matching published rounded values of 1.3% and 0.5% without target-specific fine-tuning. Persistent inference eliminated repeated model loading, while Hopper GPU execution reduced the recorded model call from 0.619 s locally to 0.202 s. The video pipeline selected 12 keyframes from a 1,674-frame stereo recording, accepted all 33 registration edges with 22 loop closures, and exported a fused model containing 16,962 points and 5,654 triangles. These results demonstrate that research-grade stereo models can be transformed into a reproducible, user-facing system for calibrated depth capture and multi-view 3D reconstruction on accessible hardware.


