To address the challenges of the large-scale georeferencing of tower-based cameras and the limited capability of video-based spatial analysis, we proposed a geospatial mapping method integrating 3D GIS and gradient descent optimization. Using a Digital Elevation Model (DEM), high-resolution remote sensing imagery, and
[...] Read more.
To address the challenges of the large-scale georeferencing of tower-based cameras and the limited capability of video-based spatial analysis, we proposed a geospatial mapping method integrating 3D GIS and gradient descent optimization. Using a Digital Elevation Model (DEM), high-resolution remote sensing imagery, and tower-based video data as the primary data sources, the proposed method first estimates the intrinsic parameters of the tower-based camera by aligning a 3D GIS virtual camera with the video imagery. Subsequently, the initial camera extrinsic parameters are estimated using the PnP algorithm based on the previously estimated intrinsic matrix
and the corresponding control point pairs. Building upon these initial estimates, the camera intrinsic and extrinsic parameters are jointly optimized using a constrained L-BFGS-B framework that incorporates prior knowledge of the tower planar location, explicit box constraints, and a semi-constrained parameterization scheme with bounded parameter ranges. Furthermore, an outlier-removal and re-optimization strategy is employed to further improve the accuracy of parameter estimation. Finally, the optimized parameters are employed to transform image coordinates into three-dimensional world coordinates, and video geospatial mapping is achieved through the integration of colored point clouds with the 3D GIS scene. The results showed the following: (1) The 3D GIS scene constructed from publicly available DEM and high-resolution remote sensing imagery met the requirements for the initial estimation of intrinsic and extrinsic camera parameters. (2) Compared with PnP, RANSAC-PnP, SQPnP, and DLT, the proposed method achieves lower reprojection and 3D spatial errors. For the independent check points, the RMSE of the reprojection error is reduced by 66.4%, 73.6%, 68.0%, and 48.3%, respectively, while the RMSE of the 3D spatial error is reduced by 84.6%, 86.2%, 83.1%, and 69.4%, respectively. These results demonstrate that the proposed method provides reliable camera parameter estimates for video geospatial mapping. (3) Using the estimated camera parameters, image coordinates are transformed into 3D world coordinates to generate a georeferenced colored point cloud, which facilitates integrated analysis with existing geospatial datasets. The proposed method provides a feasible solution for tower-based camera georeferencing and three-dimensional visualization under conditions without field calibration. It offers a theoretical and technical basis for geospatial monitoring and related applications.
Full article