Abstract
The literature on 3D point-cloud learning and road-scene geometry contains a recurring tension between methodological novelty and evidential comparability. By reading hierarchical spatial state-space aggregation for point-cloud classification alongside attention-based fusion of LiDAR and camera cues for curb detection, this article clarifies the conditions under which their conclusions can support a common research argument. A structured reading of two target studies and 12 verified companion references is conducted across five lenses: neighborhood construction, hierarchy, state-space mixing, sampling robustness, efficiency. Emphasis is placed on the provenance of evidence, the comparability of baselines, and the consequences of alternative explanations. The synthesis shows that neighborhood construction cannot be interpreted independently of hierarchy, while state-space mixing determines whether an apparent improvement remains meaningful outside the original setting. The strongest claims are therefore those that expose sensitivity, failure conditions, and residual uncertainty. The article concludes with a research agenda built around transparent comparators, targeted stress tests, and evidence records that can be reused without overstating causal or practical reach.
References
Sun, Y., Zia, A., Long, Z., Qiu, Z., Xiang, W., & Zhou, J. (2025). Hierarchical Spatial Mamba Framework for Point Cloud Classification. In Pattern Recognition and Computer Vision (pp. 402-417). Springer.
Chen, Y., Long, Z., Wu, Y., & Chen, L. (2024). Falcon: Fused attention for lidar-camera curb detection.
Zhang, T., Yuan, H., Qi, L., Zhang, J., Zhou, Q., Ji, S., et al. (2025). Point Cloud Mamba: Point Cloud Learning via State Space Model. Proceedings of the AAAI Conference on Artificial Intelligence, 39(10), 10121-10130. https://doi.org/10.1609/aaai.v39i10.33098
Tan, J., Li, J., An, X., & He, H. (2014). Robust Curb Detection with Fusion of 3D-Lidar and Camera Data. Sensors, 14(5), 9046-9073. https://doi.org/10.3390/s140509046
Song, S., Tang, K., & Zhang, Y. (2025). PST-Mamba: Spatio-temporal selective state fusion for effective point cloud video understanding with state space models. Image and Vision Computing, 163, 105785. https://doi.org/10.1016/j.imavis.2025.105785
Liu, J., Yue, S., Hao, W., & Cai, Y. (2026). MF-BEVFusion: multiscale depth estimation and fully dynamic fusion for camera-LiDAR BEV 3D object detection. Journal of Electronic Imaging, 35(02). https://doi.org/10.1117/1.jei.35.2.023009
Zhou, Z., Wang, Q., & Zhou, X. (2026). MSHI-Mamba: A Multi-Stage Hierarchical Interaction Model for 3D Point Clouds Based on Mamba. Applied Sciences, 16(3), 1189. https://doi.org/10.3390/app16031189
Bong, E. J., & Kee, S. C. (2026). Dense Depth Map Estimation Based on Camera–LiDAR Sensor Fusion. IEEE Sensors Journal, 26(4), 5891-5901. https://doi.org/10.1109/jsen.2025.3649237
Wang, G., Zhang, X., Peng, Z., Zhang, T., & Jiao, L. (2025). S 2 Mamba: A Spatial–Spectral State Space Model for Hyperspectral Image Classification. IEEE Transactions on Geoscience and Remote Sensing, 63, 1-13. https://doi.org/10.1109/tgrs.2025.3530993
Rao, R., Ouyang, Z., Chen, S., Chen, L., Huang, G., & Cui, C. (2026). Zero-Shot Polarization-Intensity Physical Fusion Monocular Depth Estimation for High Dynamic Range Scenes. Photonics, 13(3), 268. https://doi.org/10.3390/photonics13030268
Liao, J., & Wang, L. (2026). SSA-Mamba: Spatial-Spectral Attentive State Space Model for Hyperspectral Image Classification. IEEE Journal of Selected Topics in Applied Earth Observations and Remote Sensing, 19, 6403-6424. https://doi.org/10.1109/jstars.2026.3654346
Ji, M., Yang, J., & Zhang, S. (2026). DepthFusion: Depth-Aware Hybrid Feature Fusion for LiDAR-Camera 3D Object Detection. IEEE Transactions on Multimedia, 28, 7217-7227. https://doi.org/10.1109/tmm.2026.3668596
Xi, G., Wang, C., Liu, X., Xiao, B., & Wei, X. (2026). Sparse Point Cloud Classification Method Based on MSE-Mamba. Electronics, 15(14), 3087. https://doi.org/10.3390/electronics15143087
Obando-Ceron, J. S., Romero-Cano, V., & Monteiro, S. (2023). Probabilistic multi-modal depth estimation based on camera–LiDAR sensor fusion. Machine Vision and Applications, 34(5). https://doi.org/10.1007/s00138-023-01426-x
