Next Article in Journal
Dynamic Imaging Simulation and Angular Measurement Performance Degradation of Interferometric Star Trackers
Previous Article in Journal
Dual-Wavelength External Cavity Lasers Using Polymer Photonic Integrated Circuits for Optical Heterodyne RF Signal Generation
Previous Article in Special Issue
A Deep Learning-Enhanced MIMO C-OOK Scheme for Optical Camera Communication in Internet of Things Networks
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Article

An IoT-Enabled Deep Learning-Based MIMO-OFDM Scheme for Optical Camera Communication in Mobile Environments

1
Ho Chi Minh City University of Technology and Engineering, Ho Chi Minh City 700000, Vietnam
2
Institute of Research and Technology, Duy Tan University, Da Nang 550000, Vietnam
3
School of Engineering & Technology, Duy Tan University, Da Nang 550000, Vietnam
*
Author to whom correspondence should be addressed.
Photonics 2026, 13(9), 872; https://doi.org/10.3390/photonics13090872
Submission received: 7 August 2026 / Revised: 13 September 2026 / Accepted: 14 September 2026 / Published: 16 September 2026
(This article belongs to the Special Issue Optical Wireless Communications (OWC) for Internet-of-Things (IoT))

Abstract

This study proposes Multiple-Input Multiple-Output and Orthogonal Frequency-Division Multiplexing methods, as introduced in the IEEE 802.15.7a-2024 standard, to achieve higher data rates and longer transmission ranges in OCC systems where a camera is used to capture optical signals. OFDM is a multi-carrier modulation scheme extensively used in high-data-rate wireless communications to mitigate ISI caused by multipath propagation. In optical wireless communication (OWC) systems, OFDM has been widely adopted in both indoor and outdoor applications, including eHealth, smart home, and smart IoT systems. OWC technologies provide a secure and low-interference communication channel for IoT devices using visible light. In OWC-enabled edge computing, data processing is performed in nodes, reducing communication overhead and improving system scalability. Nevertheless, user mobility remains a major challenge for OWC systems, as time-varying optical channels significantly degrade signal processing performance. Furthermore, reliable signal detection under mobility is critical for improving the signal-to-noise ratio. To overcome these challenges, this paper proposes a deep learning-based LED detection scheme for a mobility-aware MIMO-OFDM system. Deep learning techniques are also utilized to identify OFDM frame boundaries and decode the transmitted data, replacing traditional signal processing approaches. Experimental results demonstrate that the proposed method enables long-range MIMO-OFDM communication over distances of up to 22 m while maintaining a low error rate at a receiver speed of 3 m/s.

1. Introduction

Due to the increasing demand for high-speed data transmission, technological advancements have been continuously pursued to enhance communication efficiency and system performance. Compared with wired systems, wireless communication provides better flexibility, easier deployment, and ubiquitous connectivity without physical connection. Therefore, a wireless communication system is a good candidate for mobile networks. However, the radio-frequency (RF) spectrum is becoming increasingly congested. Consequently, many researchers have focused on the sub-terahertz (sub-THz) band for sixth-generation (6G) mobile communications, which can potentially support data rates up to Tbps [1]. However, operation at high frequencies raises concerns regarding potential health implications. The integration of OWC with edge computing provides an effective framework for real-time, image-based wireless communication systems [2,3,4]. In [5], soft-information LLM fusion was proposed for visible light links, while massive AI models deployed on edge devices and servers were proposed in [6]. In OWC, light sources transmit data that are captured by photodiodes or image sensors. However, decoding an OWC signal requires extensive data and image processing, such as demodulation, stripe detection, and region-of-interest (RoI) detection, which might result in considerable latency when the processing is performed on a cloud server.
Visible light communication (VLC), light fidelity (LiFi), and optical camera communication (OCC) have emerged as promising alternatives to traditional RF communication systems. These optical wireless communication technologies offer several advantages. With appropriate dimming control and flicker mitigation techniques, optical communication systems can provide safe illumination and communication without posing significant health risks to humans; flicker frequencies above the critical flicker-fusion threshold are generally considered safe for human vision [7]. In addition, the optical spectrum offers substantially greater bandwidth than the RF spectrum. In addition, visible light sources are already widely deployed in existing infrastructure, such as smart homes, hospitals, and vehicles, which significantly reduces the deployment costs of VLC/LiFi/OCC technologies compared to those of RF technologies.
From this perspective, both academic institutions and industrial companies have dedicated a lot of effort to research and development in this OWC field. The IEEE 802.15 tutorials [8] provide an overview of optical wireless communication (OWC) technologies, including their protocols and system characteristics. In IEEE 802.15 standards, the IEEE 802.15.7a-2024 standard proposed a high-speed OCC technology for the Internet of Things (IoT) and vehicular applications [9]. Unlike VLC and LiFi technologies, which use photodiodes as receivers, OCC systems deploy image sensors for signal detection. Previous studies [10,11] have shown that OCC performance is strongly influenced by camera characteristics. Two types of image sensors are commonly used: global shutter cameras and rolling shutter cameras.
OFDM is a well-established and widely used modulation technique for transmitting digital information with multiple orthogonal subcarriers. It is particularly suitable for high-data-rate communications because it mitigates inter-symbol interference (ISI) by dividing the communication bandwidth into multiple narrowband subcarriers. These subcarriers can overlap in the frequency domain without interference because of their orthogonality, which is maintained through the use of the Fourier transform. To further counteract channel impairments, a cyclic prefix is appended to each OFDM symbol.
Camera on–off keying (C-OOK) was first presented in [12] as a high-speed OCC technique and was later incorporated into IEEE 802.15.7-2018. Despite its simplicity, it suffers from disadvantages such as a relatively high error rate and limited transmission range. A multiple-input multiple-output (MIMO) C-OOK system incorporating matched filtering was presented in [13] to increase data rates, though mobility effects were not addressed. However, OCC systems employing OOK are sensitive to inter-symbol interference (ISI). Subsequently, a rolling shutter OFDM (RS-OFDM) approach was introduced in [14], in which OFDM signals are reconstructed from LED intensity variations captured by rolling-shutter cameras. With a RoI algorithm for LED detection and linear equalizer, RS-OFDM has limited capability in supporting mobile environments and MIMO configurations. MIMO-OFDM was proposed in this paper by integrating MIMO technology with RS-OFDM. While this method enables high data-rate transmission, it also reveals the pronounced impact of user mobility on OCC system performance. To resolve these limitations, this paper proposes a deep learning (DL)-based LED detection model for MIMO-OFDM systems using the YOLOv11 algorithm to achieve accurate and real-time detection. In addition, a DL-based decoder is developed to improve OCC performance compared to conventional decoding methods [14]. Therefore, the proposed approach demonstrates compatibility with IEEE 802.15.7a-2024, which improves data rates, transmission distance, and mobility support compared to conventional RS-OFDM.
This paper is planned as six sections. Section 2 summarizes the technical contributions. Section 3 reviews the related MIMO-OFDM scheme for OCC technologies. Section 4 presents simulation and experimental results under mobile environments. Section 5 presents the discussion, while Section 6 concludes the paper.

2. Technical Contributions

In this study, authors present a convolutional neural network-based LED detection approach together with deep learning-based data decoding for MIMO-OFDM systems. This scheme is designed to support mobile environments for long-distance OCC considering low error rates. The main contributions of this paper are summarized as follows:
Robustness to frame-rate variations: Variations in camera frame rate represent a challenge in OCC systems. Although frame rates are often assumed to be fixed (e.g., 60 fps or 500 fps), they may fluctuate in real time due to internal camera settings, resulting in synchronization mismatches between the transmitter and receiver. The proposed approach makes data decoding reliable by using sequence numbers (SNs).
Robustness to complex noise: OCC systems are affected by various complex noise sources, including motion blur, inter-symbol interference, and optical attenuation, which are difficult to mitigate in the time domain. By applying OFDM scheme, we can mitigate these noises in the frequency domain by removing the DC component, which is difficult to eliminate in the time domain.
Mobility support: Since MIMO-OFDM systems rely on the rolling shutter effect, they are highly sensitive to mobile environments. Under the rolling shutter effect, LEDs are captured as bright and dark stripes, which limits the effectiveness of conventional region-of-interest (RoI) detection techniques in the IEEE 802.15.7-2018 standard. To address these challenges, a YOLOv11-based LED tracking algorithm is proposed to enhance robustness in mobile environments considering multiple LEDs.
Lower bit error rate: By applying MIMO technology, we can improve the bit error rate multiple times compared to the conventional OFDM scheme.
Deep learning-based decoder: Deploying a DL-based MIMO-OFDM decoder reduces the error rates under mobility conditions instead of the normal linear equalizer in [13]. In real-time optical channels, by training the proposed model on datasets collected under multiple channel conditions, multiple communication distances, multiple velocities, and multiple camera exposure times, the proposed approach achieves good performance compared to conventional methods. The details of the DL model are also explained to highlight our approach.

3. System Architecture

This section describes the architecture of the MIMO-OFDM system. In contrast to conventional OFDM technology in RF systems, where OFDM symbols are directly passed to the inverse discrete Fourier transform (IDFT), each OFDM symbol in an OCC system is processed through a Hermitian mapping block. The primary role of the Hermitian mapping is to ensure that the output of the IDFT consists of real-valued samples, as shown in Figure 1.
The definition of Discrete Fourier Transform (DFT) is shown as follows:
X [ m ] = n = 0 N 1 x [ n ] . e 2 Π . i N . m . n   with   m   =   0 ,   1 ,   2 ,   ,   N     1
In addition to that, the IDFT is represented:
x [ n ] = 1 N m = 0 N 1 X [ m ] . e 2 Π . i . m N n   with   n   =   0 ,   1 ,   2 ,   ,   N     1
The Hermitian mapping procedure is defined as follows:
X m = X N m   for   0 < m < N   and   X 0 , X 1 , X 2 X N 1
X = [ 0 , X 1 , X 2 0 X 2 , X 1 ]

3.1. Cyclic Prefix (CP)

The cyclic prefix (CP) is an important component of an OFDM symbol. It helps mitigate inter-symbol interference (ISI) in high-speed communication systems. It is created by copying the latter part of the OFDM packet and prepending it to the beginning of the symbol. Figure 2 illustrates CP creation and its insertion into the OFDM packet.

3.2. DL for LED Detection

In optical camera communication (OCC) systems, region-of-interest (RoI) [15,16,17] detection has been extensively studied, with most existing methods based on object- and feature-driven approaches that enable real-time processing. As discussed previously, the rolling shutter effect causes LEDs to appear in captured pictures as intensity-modulated stripes that represent the transmitted OFDM waveform. A single image may contain multiple such stripes, making RoI detection particularly challenging, especially in mobile scenarios.
Computer vision (CV) tasks such as object recognition, image classification, localization, and image reconstruction have widely adopted DL-based object detection models. Among these, convolutional neural networks (CNNs) have demonstrated superior performance in CV applications. The YOLO framework, which is based on a CNN architecture, is an advanced solution for real-time object detection and tracking. In addition to that, a customized YOLO model is developed and trained for LED detection and tracking, with explicit consideration of rolling shutter characteristics and mobility-induced distortions.
An experimental dataset was collected from real-world settings to validate the proposed approach. Specifically, 2500 clear and motion-blurred images were extracted using different exposure durations from video sequences that were recorded during the day and at night under mobility environments. The YOLO model was trained on NVIDIA GeForce RTX 3050 over 100 epochs. The YOLOv11 model uses a 640 × 640 input image, and a batch size of 16. The optimizer was AdamW, with a learning rate of 0.01. The dataset was split into training, validation, and test sets with a ratio of 70%, 15%, and 15%, respectively. A YOLOv11 [18,19,20,21] model with nine and eleven convolutional layers was trained using these manually annotated images. The final convolutional layer was configured with 38 filters, and a single detection class was defined.

3.3. OFDM Decoder Based on Deep Learning

The optical channel response is modeled using the Lambertian equation [22] as:
h = g o p R m cos ( i n ) d x , y   with   i n < F O V
where i n denotes the angle of incidence, g o p represents the optical filter gain, d x , y is the Euclidean range from the transmitter (Tx) to receiver (Rx), and R m is the Lambertian radiant value of the OCC channel. Under such rapidly varying conditions, the optical signal captured by image sensors may be distorted. The received optical intensity in these time-varying channels can be expressed [23] in the frequency domain.
Y [ m ] = H [ m ] X [ m ] + N [ m ]
with N is additive white Gaussian noise (AWGN), X[m] is the sent OFDM symbol, H[m] is the multipath channel response from transmitter to receiver, and Y[m] is the received OFDM waveform. The process of channel equalization is as follows:
X [ m ] = I [ m ] + Q [ m ] = Y [ m ] H [ m ]
where I[m], Q[m] are the In-phase and Quadrature parts of OFDM signal after DFT block. I [ m ] , Q [ m ] are the In-phase and Quadrature parts after channel equalization. H [ m ] is the estimated channel after channel equalization.
A deep learning-based decoder [24,25,26,27] is proposed for channel equalization. In [28], a deep learning-assisted VLC system was proposed for IoT applications, while deep learning-based channel estimation and detection were proposed in [29] for spatial modulation VLC systems. Unlike RF and VLC systems, OCC systems use image sensors for signal detection. Then we proposed the lightweight DL model for equalizer, that is suitable for OCC systems in general and MIMO-OFDM in particular, especially under mobility and long-distance transmission conditions. To evaluate its effectiveness, the proposed technology was compared with conventional algorithms, demonstrating superior performance in dynamic and mobile scenarios. The structure and operation of the proposed DL model for channel estimation and equalization are illustrated in Figure 3. The signal after the DFT block is fed into the input of the DL model. By training the model on a dataset containing both clear and motion-blurred images across various transmission distances, data can be effectively decoded, as shown in Figure 3. The output represents the predicted response of the channel equalizer.
Channel equalization for mobile environments using deep learning, illustrated in Figure 3, utilizes only two hidden layers and two input/output layers to minimize model complexity and prevent overfitting in models, with four neurons in each hidden layer. The ReLU activation function was applied to the hidden layers. The mean absolute error (MAE) was employed as the loss function, and the Adam optimizer was used for training, with a learning rate of 0.001. The dataset was divided to use 80% of it for training and 20% for validation, with a batch size of 32. The model was trained on an NVIDIA GeForce RTX 3050 for 30 epochs. The network accurately estimated the in-phase and quadrature components of OFDM signal, resulting in superior decoding performance compared to conventional methods in mobile environments. During the training process, increasing the number of hidden layers beyond five led to a case of overfitting, reducing performance. The proposed approach performed well in laboratory environments, with communication distances ranging from 2 to 22 m and mobility speeds of 0–3 m/s.

4. Simulation and Implementation Results

4.1. BER Estimation for O-OFDM Technologies

With optical OFDM, a symmetric clipping operation is required for DCO-OFDM systems [30]. In addition, ACO-OFDM can be employed to mitigate errors caused by bias clipping at the upper limit of the OFDM signal. In this study, K denotes the attenuation factor applied during the clipping process, while β b o t t o m , β t o p represent the upper and lower clipping thresholds, respectively. The formulation follows the expressions presented in [31,32]:
K = C o v ( s , ψ ( s ) ) σ 2 = Q ( β b o t t o m ) Q ( β t o p )
with Cov[.] representing the covariance operator, σ 2 representing the signal’s variance, and ψ ( . ) showing the normalized nonlinear transfer function.
The BER values of M-QAM O-OFDM are represented as [32]:
B E R = 4 ( M 1 ) M log 2 ( M ) Q ( 3 log 2 ( M ) M 1 Γ b ( e l e c ) ) + 4 ( M 2 ) M . log 2 ( M ) Q ( 3 3 log 2 ( M ) M 1 Γ b ( e l e c ) )
where M is the symbol QAM’s number constellation in OFDM. Γ b ( e l e c ) stands for received electrical SNR per bit on enabled subcarriers in M-QAM DCO-OFDM. Γ b ( e l e c ) is distinguished as follows:
Γ b ( e l e c ) = K 2 P b ( e l e c ) / G B σ c l i p 2 + G B σ A W G N 2 g h ( o p t ) 2 G D C = K 2 G B σ c l i p 2 P b ( e l e c ) + G B γ b ( e l e c ) 1 g h ( o p t ) 2 G D C
where γ b ( e l e c ) = E b / N 0 denotes the electrical SNR in each bit; σ c l i p 2 represents the variance of the clipping noise; G D C is the attenuation factor of the electrical signal power; g h ( o p t ) denotes the optical path gain coefficient; G B is the utilization ratio of subcarriers in DCO-OFDM with G = N 2 N ; N denotes the OFDM symbol. In addition, for ACO-OFDM, G B = 0.5
Figure 4 and Figure 5 illustrate the simulation of BER values of ACO and DCO-OFDM as a function of pixel E b / N 0 , considering appropriate upper and lower clipping stages for the parameters used O-OFDM. The results are obtained in Figure 4 and Figure 5 by Monte Carlo simulations with 100 independent runs. The results indicate that, to achieve a target BER of 10 4 , the required pixel E b / N 0 has to be at least 22 dB for DCO-OFDM and 17 dB for ACO-OFDM. As discussed earlier, pixel SNR can improve by employing higher exposure times; however, this comes at the expense of reduced communication bandwidth. Consequently, careful control of the communication bandwidth is necessary [14]. From the results in Figure 4 and Figure 5, we can easily design an OCC system with some parameters (LED power, distance between two LEDs, camera parameters) for real-world environments to achieve the target BER as in the simulation results.

4.2. Implementation

Frame rate variations in cameras introduce significant challenges to OCC systems. They lead to oversampling and undersampling cases. When the camera frame rate is more than two times higher than the packet rate, then we can get similar images in the data collecting process. To resolve this issue, a sequence number (SN) is embedded within the data structure (DS) to manage frame rate variations. By evaluating the SNs of the received DSs, the receiver can identify and discard duplicate packets; as illustrated in Figure 6a, it discards packets with identical consecutive SNs and merges sequential SNs (n, n + 1, n + 2).
Conversely, the undersampling case occurs when the frame rate drops below the packet rate, leading to lost payloads during processing. Figure 6b demonstrates how missing packets are detected using SNs. The SN bit length determines the maximum number of consecutive lost packets that can be detected. For example, a 2-bit SN can detect up to four missing payloads by detecting non-sequential transitions, such as from n to n + 2, as shown in Figure 6b.
To examine the effects of frame rate variations, the MIMO-OFDM system was evaluated several times using various cameras. The sequence number (SN) length was suitably chosen for the asynchronous processing method. A PointGrey camera with a variable frame rate ranging from 40 to 60 frames per second (fps) was used to evaluate the MIMO-OFDM system. Figure 7 displays the received MIMO-OFDM waveforms captured by the camera. An OFDM size of 128 or 256 FFT points was deployed with a 20% cyclic prefix (CP). A PointGrey rolling shutter camera operating at 60 fps with a 35 mm focal-length lens was used to receive MIMO-OFDM signals from two LEDs (10 V DC-2.5 W) at different distances. The proposed OCC system achieved a latency of less than 30 milliseconds, including the frame-gap time and decoder processing time. The DCO-OFDM was chose in our implementation instead of ACO-OFDM due to the higher data rate of DCO-OFDM with FFT sizes of 128 or 256 bits.
The desired sequence numbers (SNs) were incorporated into the asynchronous packets containing the transmitted data. To evaluate system performance, the experiments also assessed different optical clock rates for two frames sizes: 128 bits and 256 bits. The findings show that a frame length of 128 bits achieves a data rate of 5.12 kbps, whereas a length of 256 bits results in a data rate of 10.24 kbps. As discussed earlier, although higher data rates can be attained with longer packet lengths, the trade-off between the LED size and the achievable data rate should be cautiously considered [12].
Figure 8 shows the MIMO-OFDM performance with non-deep learning and deep learning at various transmission distances (2–22 m) with an exposure time of 70 μs at a speed of 3 m/s. The dataset collected data not only at 3 m/s but also at velocities ranging from 0 to 3 m/s. Therefore, the DL model was expected to support velocities of up to 3 m/s. We selected 3 m/s for Figure 7 because it represented the most challenging mobility condition considered in our experiments. The experimental setup was configured as follows: the distance between the two LEDs was 2 cm, and the LEDs and camera were positioned 1 m above the ground. The transmitter and receiver were aligned to establish a line-of-sight (LoS) link. During the experiments, the receiver was moved back and forth along the transmission direction. To validate the proposed approach, an experimental dataset was collected from real-world settings, comprising 2500 clear and motion-blurred images extracted under different exposure times from video sequences recorded during both daytime and nighttime under mobile conditions. A YOLOv11 model was trained using these manually annotated images. The proposed YOLOv11 model achieved high detection performance (precision of 0.95, recall of ~0.96, F1-score of ~0.95, and m A P 50 95 of ~88.1).
To measure the bit error rate, we transmitted 1,000,000 packets, with each packet containing 128 or 256 bits. The transmission experiment was repeated at least 10 times over same distances with DL model. Using the DL decoder, a BER of 10 5 was achieved at a distance of 2 m, whereas the conventional approach attained only 10 4 . As the transmission distance increases, the received optical intensity from the LED decreases, resulting in a lower SNR. Since a lower SNR leads to a higher BER, the BER is expected to increase with the increasing transmission distance. Under identical distance and environmental conditions, the DL-based decoder improved OCC performance in mobility more than the non-DL method, which used the RoI algorithm for LED detection and a linear equalizer for channel estimation. In addition, DL was employed for LED tracking, which further supports accurate data decoding and enhances performance at longer communication distances and under mobile conditions. In the implementation, a movement was adopted to ensure suitability for indoor applications while accounting for mobility effects. In this study, we deployed YOLOv11 for LEDs detection, achieving precision of over 95% considering the two LEDs at 3 m/s speed. To verify the effect of mobility conditions under indoor environments, we tested at 3 m/s, which is suitable for the walking speed. To enhance OCC performance, it is possible to collect and train comprehensive datasets encompassing various mobile environments, ensuring that the deep learning model adapts effectively to different mobility conditions. The Supplementary Materials section shows the proposed approach at a distance of 1 m, with visual illustrations provided in the supplementary documentation of this study.

5. Discussion

The rolling shutter effect causes LEDs to appear in captured images as stripes that represent the transmitted OFDM signals. A captured image may include multiple stripes, making RoI detection particularly challenging, especially in mobile scenarios. Computer vision (CV) algorithms, such as object recognition, image classification, localization, and image reconstruction, have made use of DL-based object detection. Among these models, convolutional neural networks (CNNs) have demonstrated excellent performance in various CV applications. The YOLO framework, based on CNN architectures, is a state-of-the-art algorithm for real-time object detection and tracking. Therefore, customized YOLO models were developed and trained for LED detection and tracking, with explicit consideration of rolling shutter characteristics and mobility-induced distortions. The proposed MIMO-OFDM scheme demonstrates that combining the rolling-shutter effect with a deep learning-based decoder significantly improves communication robustness under receiver mobility. As previously mentioned, real mobile scenarios are characterized by distortion, motion blur, and intermittent streaking. With the proposed DL-based decoder, reliable data recovery is achieved even when conventional decoding techniques fail. These results indicate that DL can effectively compensate for moving environments and long distances. The experimental results also demonstrate that the proposed approach improves BER performance under mobile conditions. This confirms that the DL model can address the challenges of mobile environments in MIMO-OFDM system. To enhance OCC performance, it is possible to collect and train comprehensive datasets encompassing various mobile environments, ensuring that the deep learning model adapts effectively to any mobile conditions. The BER results in Figure 8 show the advantage of the DL model compared to conventional methods. The proposed scheme is also compatible with the IEEE 802.15.7a-2024 standard for higher-rate and longer-range OCC, since the neural network operates as a preprocessing step before OFDM demodulation. This architecture provides the flexibility to integrate DL for object detection and DL for decoding with a MIMO-OFDM scheme in OCC systems. Overall, this study highlights the possibilities of deep learning models used in OCC systems. By complementing traditional methods, DL increases OCC performance under mobile conditions. The limitations of this study are that the proposed system was evaluated only at velocities of 0–3 m/s and distances of 2–22 m. At longer distances and higher velocities, the proposed system did not perform well. In future work, we will collect a more comprehensive dataset covering higher velocities and longer distances. In addition, we will improve the deep learning models to make them more robust and suitable for harsh operating environments.

6. Conclusions

This paper uses deep learning techniques to develop a MIMO-OFDM scheme for mobile environments. Deep learning techniques are used for data decoding as well as LED detection and tracking. To identify several LEDs, DL is applied. Accurate LED recognition is more challenging than traditional RoI-based techniques because of the rolling shutter effect, which causes LEDs to appear in collected images as alternating light and dark strips. Furthermore, long-distance communication and mobile environments performances are improved by a deep learning-based decoder. The proposed deep learning-based MIMO-OFDM scheme’s bit error rate performance is evaluated at different distances and compared with a non-deep learning approach. The results show that the MIMO-OFDM method based on DL significantly reduces BER under mobility conditions.

Supplementary Materials

The Supplementary Materials have been uploaded at: https://zenodo.org/records/19449409 (accessed on 7 April 2026).

Author Contributions

Conceptualization, V.K.P.; Methodology, V.K.P.; Software, V.K.P.; Formal analysis, H.N.; Writing—original draft, H.N.; Visualization, H.N.; Supervision, H.N.; Project administration, V.K.P.; Funding acquisition, H.N. All authors have read and agreed to the published version of the manuscript.

Funding

This research received no external funding.

Data Availability Statement

Restrictions apply to the datasets due to the project policy.

Conflicts of Interest

The authors declare no conflicts of interest.

References

  1. Pan, Z.; Xu, Z.; Miao, R.; Zhao, T.; Wang, J. Prospects of 6G Technology Framework: A Big-Lite Multi-RATs Concept. IEEE Commun. Mag. 2025, 63, 174–180. [Google Scholar] [CrossRef] [Scilit]
  2. Khoshafa, M.H.; Maraqa, O.; Moualeu, J.M.; Aboagye, S.; Ngatched, T.M.; Ahmed, M.H.; Gadallah, Y.; Di Renzo, M. RIS-Assisted Physical Layer Security in Emerging RF and Optical Wireless Communications Systems: A Comprehensive Survey. IEEE Commun. Surv. Tutor. 2025, 27, 2156–2203. [Google Scholar] [CrossRef] [Scilit]
  3. Chia, L.W.; Motani, M. High-Performance OCC with Edge Processing on SPAD and Event-Based Cameras. IEEE Commun. Mag. 2024, 62, 62–67. [Google Scholar] [CrossRef] [Scilit]
  4. Bhutani, M.; Lall, B.; Agrawal, M. Optical Wireless Communications: Research Challenges for MAC Layer. IEEE Access 2022, 10, 126969–126989. [Google Scholar] [CrossRef] [Scilit]
  5. Thai, P.Q. Soft-Information LLM Fusion for LDPC-Coded Text Over Visible-Light Links. IEEE Commun. Lett. 2026, 30, 1885–1889. [Google Scholar] [CrossRef] [Scilit]
  6. Lyu, Z.; Xiao, M.; Xu, J.; Skoglund, M.; Di Renzo, M. The Larger the Merrier? Efficient Large AI Model Inference in Wireless Edge Networks. IEEE J. Sel. Areas Commun. 2026, 44, 2839–2853. [Google Scholar] [CrossRef] [Scilit]
  7. Sridhar, R.; Richard, D.; Kyu, L.S. IEEE 802.15.7 Visible Light Communication: Modulation and Dimming Support. IEEE Commun. Mag. 2012, 50, 72–82. [Google Scholar] [CrossRef] [Scilit]
  8. Nikola, S.; Volker, J.; Min, J.Y.; John, L.Q. An Overview on High-Speed Optical Wireless/Light Communications. 2017. Available online: https://mentor.ieee.org/802.11/dcn/17/11-17-0962-02-00lc-an-overview-on-high-speed-optical-wireless-light-communications.pdf (accessed on 13 January 2026).
  9. IEEE-SA. IEEE Standard for Local and Metropolitan Area Networks—Part 15.7: Short-Range Optical Wireless Communications Amendment 1: Higher Rate, Longer Range Optical Camera Communication (OCC); IEEE-SA: Piscataway, NJ, USA, 2024. [Google Scholar]
  10. Nguyen, H.; Al-Imran; Jang, Y.M. Survey of next-generation optical wireless communication technologies for 6G and Beyond 6G. ICT Express 2025, 11, 576–589. [Google Scholar] [CrossRef] [Scilit]
  11. Dong, K.; Kong, M.; Wang, M. Error performance analysis for OOK modulated optical camera communication systems. Opt. Commun. 2025, 574, 131121. [Google Scholar] [CrossRef] [Scilit]
  12. Std 802.15.7-2018; IEEE Standard for Local and Metropolitan Area Networks—Part 15.7: Short-Range Optical Wireless Communications. IEEE-SA: Piscataway, NJ, USA, 2018.
  13. Nguyen, D.T.; Nguyen, T.; Thieu, M.D.; Nguyen, H. A Deep Learning-Enhanced MIMO C-OOK Scheme for Optical Camera Communication in Internet of Things Networks. Photonics 2026, 13, 163. [Google Scholar] [CrossRef] [Scilit]
  14. Nguyen, H.; Thieu, M.D.; Nguyen, T.; Jang, Y.M. Rolling OFDM for Image Sensor Based Optical Wireless Communication. IEEE Photonics J. 2019, 11, 6500817. [Google Scholar] [CrossRef] [Scilit]
  15. Yu, Q.; Wang, B.; Su, Y. Object Detection-Tracking Algorithm for Unmanned Surface Vehicles Based on a Radar-Photoelectric System. IEEE Access 2021, 9, 57529–57541. [Google Scholar] [CrossRef] [Scilit]
  16. Lin, H.; Si, J.; Abousleman, G.P. Region-of-interest detection and its application to image segmentation and compression. In Proceedings of the 2007 International Conference on Integration of Knowledge Intensive Multi-Agent Systems, Waltham, MA, USA, 30 April–3 May 2007. [Google Scholar]
  17. Yan, C.; Chen, W.; Chen, P.C.Y.; Kendrick, A.S.; Wu, X. A new two-stage object detection network without RoI-Pooling. In Proceedings of the 2018 Chinese Control and Decision Conference (CCDC), Shenyang, China, 9–11 June 2018; pp. 1680–1685. [Google Scholar]
  18. Zhang, H.; Gao, L.; Gong, Y.; Liu, H.; Zhu, Y.; Yang, Y. RTF-SAW-YOLOv11: A Bolt Defect Detection Model for Power Transmission Lines Under Low-Light Conditions. IEEE Access 2025, 13, 138640–138659. [Google Scholar] [CrossRef] [Scilit]
  19. Luo, C.; Tang, H.; Li, S.; Wan, G.; Chen, W.; Guan, J. YOLOv11s-CD: An Improved YOLOv11s Method for Catenary Dropper Fault Detection. IEEE Trans. Instrum. Meas. 2025, 74, 5043410. [Google Scholar] [CrossRef] [Scilit]
  20. Xue, Z.; Kong, L.; Wu, H.; Chen, J. Fire and Smoke Detection Based on Improved YOLOV11. IEEE Access 2025, 13, 73022–73040. [Google Scholar] [CrossRef] [Scilit]
  21. Zhang, L.; Zheng, A.; Sun, X.; Sun, Z. Enhanced YOLOv11-Based River Aerial Image Detection Research. IEEE Geosci. Remote Sens. Lett. 2025, 22, 8002405. [Google Scholar] [CrossRef] [Scilit]
  22. Ghassemlooy, Z.; Alves, L.N.; Zvanovec, S.; Khalighi, M.A. Visible Light Communications: Theory and Applications, 1st ed.; CRC Press: Boca Raton, FL, USA, 2016. [Google Scholar]
  23. Aziz, M.A.; Rahman, M.H.; Sejan, M.A.S.; Tabassum, R.; Hwang, D.D.; Song, H.K. Deep Recurrent Neural Network Based Detector for OFDM With Index Modulation. IEEE Access 2024, 12, 89538–89547. [Google Scholar] [CrossRef] [Scilit]
  24. Lee, H.; Lee, S.H.; Quek, T.Q.S.; Lee, I. Deep Learning Framework for Wireless Systems: Applications to Optical 384 Wireless Communications. IEEE Commun. Mag. 2019, 57, 35–41. [Google Scholar] [CrossRef] [Scilit]
  25. Wu, H.; Chen, Z.; Liu, Z.; Geng, X.; Zhao, Y.; Liu, Z. CRS-Based Joint CFO and Channel Estimation Using Deep Learning in OFDM-Based Vehicular Communication Systems. IEEE Trans. Wirel. Commun. 2025, 24, 3882–3893. [Google Scholar] [CrossRef] [Scilit]
  26. Kong, M.; Pan, Y.; Zhou, H.; Yu, R.; Le, X.; Yuan, H.; Wang, R.; Yang, Q. Deep Learning-Based Acquisition Pointing and Tracking for Underwater Wireless Optical Communication. IEEE Photonics Technol. Lett. 2025, 37, 555–558. [Google Scholar] [CrossRef] [Scilit]
  27. Jia, B.; Ge, W.; Cheng, J.; Du, Z.; Wang, R.; Song, G.; Zhang, Y.; Cai, C.; Qin, S.; Xu, J. Deep Learning-Based Cascaded Light Source Detection for Link Alignment in Underwater Wireless Optical Communication. IEEE Photonics J. 2024, 16, 7801512. [Google Scholar] [CrossRef] [Scilit]
  28. El Jbari, M.; Ettehamy, Z.; Moussaoui, M.; Menhaj, A.R.; de Figueiredo, F.A.; Ouameur, M.A. Deep learning-assisted intelligent VLC for IoT applications: A systematic review towards digital and autonomous 6G wireless networks. Sci. Afr. 2026, 33, e03556. [Google Scholar] [CrossRef] [Scilit]
  29. Palitharathna, K.W.S.; Suraweera, H.A.; Godaliyadda, R.I.; Herath, V.R.; Thompson, J.S. Neural Network-Based Channel Estimation and Detection in Spatial Modulation VLC Systems. IEEE Commun. Lett. 2022, 26, 1598–1602. [Google Scholar] [CrossRef] [Scilit]
  30. Randel, S.; Breyer, F.; Lee, S.C.; Walewski, J.W. Advanced modulation schemes for short-range optical communications. IEEE J. Sel. Top. Quantum Electron. 2010, 16, 1280–1289. [Google Scholar] [CrossRef] [Scilit]
  31. Zhou, J.; Wang, Q.; Cheng, Q.; Guo, M.; Lu, Y.; Yang, A.; Qiao, Y. Low-PAPR Layered/ Enhanced ACO-SCFDM for Optical Wireless Communications. IEEE Photonics Technol. Lett. 2018, 30, 165–168. [Google Scholar] [CrossRef] [Scilit]
  32. Dimitrov, S.; Sinanovic, S.; Haas, H. Clipping noise in OFDM based Optical Wireless Communication Systems. IEEE Trans. Commun. 2012, 60, 1072–1081. [Google Scholar] [CrossRef] [Scilit]
Figure 1. System architecture of MIMO-OFDM based on deep learning.
Figure 1. System architecture of MIMO-OFDM based on deep learning.
Photonics 13 00872 g001
Figure 2. The OFDM symbol’s cyclic prefix.
Figure 2. The OFDM symbol’s cyclic prefix.
Photonics 13 00872 g002
Figure 3. Channel equalization for mobility environments using deep learning.
Figure 3. Channel equalization for mobility environments using deep learning.
Photonics 13 00872 g003
Figure 4. ACO-OFDM simulation in relation to pixel SNR.
Figure 4. ACO-OFDM simulation in relation to pixel SNR.
Photonics 13 00872 g004
Figure 5. BER DCO-OFDM simulation in relation to pixel SNR.
Figure 5. BER DCO-OFDM simulation in relation to pixel SNR.
Photonics 13 00872 g005
Figure 6. (a) Oversampling case for the merge packet algorithm; (b) undersampling case for missing packet detection.
Figure 6. (a) Oversampling case for the merge packet algorithm; (b) undersampling case for missing packet detection.
Photonics 13 00872 g006
Figure 7. Rx Interface.
Figure 7. Rx Interface.
Photonics 13 00872 g007
Figure 8. Performance of the MIMO-OFDM scheme at a velocity of 3 m/s.
Figure 8. Performance of the MIMO-OFDM scheme at a velocity of 3 m/s.
Photonics 13 00872 g008
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Pham, V.K.; Nguyen, H. An IoT-Enabled Deep Learning-Based MIMO-OFDM Scheme for Optical Camera Communication in Mobile Environments. Photonics 2026, 13, 872. https://doi.org/10.3390/photonics13090872

AMA Style

Pham VK, Nguyen H. An IoT-Enabled Deep Learning-Based MIMO-OFDM Scheme for Optical Camera Communication in Mobile Environments. Photonics. 2026; 13(9):872. https://doi.org/10.3390/photonics13090872

Chicago/Turabian Style

Pham, Van Khoa, and Huy Nguyen. 2026. "An IoT-Enabled Deep Learning-Based MIMO-OFDM Scheme for Optical Camera Communication in Mobile Environments" Photonics 13, no. 9: 872. https://doi.org/10.3390/photonics13090872

APA Style

Pham, V. K., & Nguyen, H. (2026). An IoT-Enabled Deep Learning-Based MIMO-OFDM Scheme for Optical Camera Communication in Mobile Environments. Photonics, 13(9), 872. https://doi.org/10.3390/photonics13090872

Note that from the first issue of 2016, this journal uses article numbers instead of page numbers. See further details here.

Article Metrics

Back to TopTop