LEARNet: A Learning Entropy-Aware Representation Network for Educational Video Understanding
Abstract
1. Introduction
- A Novel Formulation of Instructional Structure: We identify an Instructional Steady-State in educational videos, characterized by long periods of visual consistency interrupted by brief, semantically rich transitions. Capturing this structure requires selecting frames that preserve pedagogically significant content, rather than relying on low-level change detection.
- The Domain-Specific LEARNet Framework: We propose LEARNet, a deep learning architecture tailored to educational videos that integrates two novel components:
- Temporal Information Bottleneck (TIB): Filters redundant frames using multi-modal cues to detect meaningful transitions and preserves only pedagogical significant frames.
- Spatial-Semantic Decoder (SSD): Annotates these keyframes by constructing high-fidelity semantic graphs, capturing both the spatial layout and the semantic relationships among educational components, such as slides, diagrams, formulas, and titles. The Relational Consistency Verification Network (RCVN) ensures that these semantic relationships are coherent, minimizing errors and preserving instructional meaning.
- EVUD-2M Benchmark: We introduce EVUD-2M, a two-million-frame educational video dataset generated by LEARNet. Starting from a 200-frame reference set, LEARNet applied a one-shot, semi-automatic annotation pipeline to produce fine-grained, multi-tier annotations for slides, diagrams, equations, and handwritten text. EVUD-2M integrates diverse sources and instructional formats, providing a scalable, robust benchmark for hierarchical educational video understanding and demonstrating LEARNet’s effectiveness as a domain-specific, entropy-aware annotation framework.

2. Literature Survey
2.1. Keyframe Extraction Techniques: A Paradigm Ill-Suited for Educational Content
2.1.1. High-Entropy Feature Methods
2.1.2. Mid-Level Feature Methods
2.1.3. Semantic and the Information-Theoretic Methods
2.2. Semantic Annotation: A Fragmented Landscape Lacking Educational Focus
2.2.1. Specialized but Static Datasets
2.2.2. Coarse Temporal Segmentations
2.2.3. General-Purpose Multimodal Models
2.2.4. The Gap in Scalable, Finely Annotated Benchmarks
2.2.5. The Integration Gap in Foundational Models
- Rely on flawed temporal assumptions, emphasizing high-entropy cues like motion rather than detecting low-entropy but pedagogically significant changes, which leads to inefficient compression of instructional information.
- Annotate at inappropriate granularities, resulting in incomplete semantic decoding that misses concentrated, high-value educational content.
- Do not provide a domain-specific framework capable of optimizing information retention, leaving the potential of modern models underutilized.
3. The EVUD-2M Benchmark: A Dataset for Hierarchical Educational Video Parsing
3.1. Curation Strategy and Semantic Coverage
- Authentic Lecture Context: The benchmark draws primarily from the NPTEL and ClassX repositories, offering a broad and realistic collection of lecture videos. This ensures models are exposed to the inherent challenges of the Instructional Steady-State, including static layouts, subtle transitions, and a mix of visual aids.
- Preventing Layout Bias: To avoid overfitting to specific video styles, a large set of standalone slides from SlideShare-1M is included. This helps models focus on understanding the actual meaning within slides, diagrams, and tables, instead of depending on video layout or recording patterns.
- Targeted Component Enhancement: Additional data on figures and tables is incorporated to strengthen model learning on key visual instructional elements. This enrichment supports fine-grained, multi-label annotations and enables precise scene parsing aligned with the goals of the LEARNet framework.
3.2. Dataset Curation and Composition
| Dataset Name | Content Type | Key Statistics | Usage in Work |
|---|---|---|---|
| NPTEL [42] | Educational lecture videos |
| Primary source for authentic, diverse lecture content across STEM and humanities. |
| ClassX [50] | Lecture video clips |
| Supplements with varied lecture styles and institutional presentation formats. |
| SlideShare-1M [61] | Presentation slides |
| Provides high-quality, standalone slides to prevent model overfitting to specific video layouts. |
| Tablenet [62], TableBank [56] | Scanned document images |
| Enables robust detection of specific educational components (tables, plots, diagrams) through targeted training |
| Roboflow Platform [63] | Computer vision datasets |
|
3.3. Comparison with State-of-the-Art Educational Video Datasets
- Intelligent frame selection: Instead of sampling frames at fixed intervals, we employ a pedagogical significance scorer that identifies the most instructionally critical moments.
- Educational Element Annotation: The benchmark provides detailed, region-level labels for academic components like equations and diagrams, moving beyond generic object detection.
- Inherent Generalization: By leveraging modern vision-language models, the dataset equips AI systems with zero-shot capabilities, allowing them to recognize educational concepts they were not explicitly trained on.
- Unified multi-format Analysis: It seamlessly handles the full spectrum of instruction—from slides and digital whiteboards to traditional blackboards—within a single, coherent framework.
3.4. Data Preprocessing for Robust Parsing
- Frame Normalization and Quality Enhancement: Video frames were standardized and resized to the resolutions required by the downstream modules. Adaptive filtering was applied to reduce compression artifacts and enhance clarity without altering pedagogical meaning.
- Illumination Normalization: Histogram equalization and gamma correction were used to stabilize brightness and contrast, improving robustness to varied classroom lighting and recording setups.
- Semantic-Preserving Augmentation: Light augmentations such as mild rotations, flips, and subtle color adjustments were introduced to improve generalization. Only transformations that preserved the semantic interpretation of diagrams, text regions, and other instructional elements were applied.
3.5. Multilevel Annotation Framework
- Image-Level Filtering: Frames are first categorized to separate pedagogically relevant scenes (e.g., slides, board content) from non-informative visuals (e.g., blank screens, logos), ensuring computational effort is focused on meaningful instructional content.
- Region-Level Spatial Annotation: Pixel-accurate segmentation masks are provided for nine educational element categories—including text blocks, diagrams, equations, and titles—supporting downstream tasks such as OCR, diagram parsing, and fine-grained content retrieval.
- Relational-Level Semantic Graphing: Triplet-based annotations (e.g., <figure, has, figure_title>) model relationships between components, capturing instructional structure and linking visual elements with their explanatory context.
3.6. Annotation Scaling and Verification
4. Methodology
4.1. Overall Method
- Structural Consistency Change to detect scene transitions,
- Domain-Weighted Saliency to emphasize educationally important regions, and
- Semantic Cluster Coherence to ensure each keyframe represents a distinct instructional unit.
| Algorithm 1: Temporal Information Bottleneck (TIB) |
|
| Algorithm 2: Spatial-Semantic Decoder (SSD) |
|
4.2. Temporal Information Bottleneck (TIB)
- Structural dissimilarity: The model first quantifies macro-level scene changes through pixel-wise dissimilarity analysis between consecutive frames. This captures significant visual transitions such as slide changes or scene shifts, serving as the foundational signal for major content boundaries. This represents a coarse measure of signal-level change. The structural consistency between two frames is calculated using MSE as given in Equation (4):
- 2.
- Content Saliency Analysis: To address the limitation of structural metrics in detecting subtle pedagogical changes, we incorporate content-aware saliency mapping. Identifies regions of visual importance within each frame, crucial for capturing incremental content additions (e.g., handwritten equations) that may not trigger significant pixel-level changes but constitute new information. The saliency score for a pixel in frame can be computed as in Equation (5):
- 3.
- Pedagogical Information Score (PIS): The structural and saliency components are integrated through an adaptive weighting scheme to compute the unified Pedagogical Information Score (Algorithm 1, Line 6). This is used in conjunction with the ISS principle to identify instructionally stable frames. The weight parameter α = 0.67, optimized empirically, balances the contribution of broad structural changes against focused content saliency, ensuring robust performance across diverse educational video formats. Unlike conventional saliency fusion, Equation (6) derives its adaptive weight α from the entropy of frame-level instructional cues, aligning selection with the information bottleneck principle.
- 4.
- Instructional Transition Identification: The temporal sequence of Pedagogical Scores undergoes Density Peak Clustering (DPC) (Algorithm 1, Line 9) to identify the most significant and non-redundant frames. This clustering approach operates on the principle that instructionally important frames form natural density peaks in the feature space, with cluster centroids representing optimal keyframe candidates.
- Temporal Significance Filters: Leveraging frame duration metadata, we prioritize frames with substantial exposition time while discarding transient content:
- Extended duration Frames: Frames with extended screen time (Equation (9)) are prioritized as they often contain stable, pedagogically rich content.
- Brief segments: Frames with very brief screen time or short associated transcription length are considered transient and are discarded as in Equation (10).
- Content Quality validation: We eliminate non-educational elements through template matching.
- Blank Page and Logo Detection: Frames matching pre-defined templates for blank slides or institutional logos are detected using Normalized Cross-Correlation (NCC) and discarded to avoid processing non-content frames as in Equation (11).NCC for images A and B is computed as in Equation (12):
- Instructional Context Verification: As mentioned in Equation (13), frames dominated by a human speaker without accompanying educational content (e.g., slides, text) are removed. A frame is discarded if the area of detected persons exceeds a threshold and a context-checking function (e.g., via object detection for slides/text) returns false.
- Temporal Coherence Optimization: This step leverages the Structural Similarity Index Measure (SSIM) maintain semantic diversity while reducing redundancy.
- High-similarity consecutive frames: Frames highly like their immediate predecessors are considered redundant and pruned as in Equation (14).
- Moderate-Significance Bucketing: Frames that are moderately like previous keyframes and possess low visual saliency are deemed of secondary importance. These are not discarded but are stored in a separate set as in Equation (15) for potential use in applications requiring higher temporal density.
Parameter Optimization and Sensitivity Analysis
4.3. Spatial Semantic Decoder (SSD)
- Pedagogical ambiguity resolution (distinguishing diagrams, equations, and complex illustrations)
- Structural diversity management (consistent labeling across slides, blackboards, and digital writing)
- Temporal consistency maintenance across evolving lecture segments
4.3.1. Model Optimization
4.3.2. Loss Function and Learning Strategy
4.3.3. Computational Efficiency and Trade-Offs
4.4. Integrated Framework and EVUD-2M Benchmark Construction
- Temporal pedagogical coherence through carefully selected keyframes
- Spatial semantic richness with region-level annotations for educational elements
- Structural diversity covering slides, diagrams, equations, and handwritten content
4.5. Implementation and Experimental Framework
5. Results and Analysis
5.1. Evaluation Framework and Metrics
5.2. Comparative Baselines: Establishing the State of the Art
5.2.1. Keyframe Extraction Baselines
5.2.2. Semantic Annotation Baselines
5.2.3. Specialized Educational Systems
5.3. LEARNet Performance Analysis
5.3.1. Optimization of Pedagogical Information Score (PIS)
5.3.2. Value Basis and Sensitivity of the DPC Cut-Off Parameter
5.3.3. TIB Performance for Pedagogical Keyframe Extraction
5.3.4. Comparison of Entropy Reduction with Traditional Non-Computational Methods
5.3.5. SSD Performance on Semantic Annotation
Comparative Analysis with Domain-Specific Educational Models
Cross Format Validation and Generalization
Per-Category Performance Analysis
Ablation Studies: Validating the Integrated Architecture
5.3.6. Specialization Advantage over Foundation Models
5.4. Evaluation on Robustness and Generalization of LEARNet
- Low Resolution: down sampling and up sampling to simulate pixel-level loss.
- Compression Artifacts: controlled JPEG degradation at varying quality factors.
- Motion Blur/Jitter: Gaussian linear blur simulating camera shake.
6. Discussion
6.1. Performance Analysis: Advancing Educational Video Parsing
6.2. The EVUD-2M Benchmark: Establishing a New Standard for Semantically Rich Educational Video Data
6.3. The Impact of Educational Structure Modeling
6.4. Cross-Format Performance Analysis
- Slides: 0.895 F1—clean layouts and stable formatting support high accuracy.
- Blackboard content: 0.879 F1—performance remains strong despite handwriting variability.
- Digital writing: 0.725 F1—lower performance due to transient strokes, motion noise, and limited structural regularity.
6.5. Limitations and Future Work
7. Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Koprinska, I.; Carrato, S. Temporal Video Segmentation: A Survey. Signal Process Image Commun. 2001, 16, 477–500. [Google Scholar] [CrossRef]
- Smeulders, A.W.M.; Worring, M.; Santini, S.; Gupta, A.; Jain, R. Content-Based Image Retrieval at the End of the Early Years. IEEE Trans. Pattern Anal. Mach. Intell. 2000, 22, 1349–1380. [Google Scholar] [CrossRef]
- Islam, R.; Moushi, O.M. GPT-4o: The Cutting-Edge Advancement in Multimodal LLM. In Intelligent Computing. Proceedings of the 2025 Computing Conference; Springer: Berlin/Heidelberg, Germany, 2024; Volume 4. [Google Scholar]
- Chen, B.-W.; Wang, J.-C.; Wang, J.-F. A Novel Video Summarization Based on Mining the Story-Structure and Semantic Relations Among Concept Entities. IEEE Trans. Multimed. 2009, 11, 295–312. [Google Scholar] [CrossRef]
- Zhao, H.; Wang, W.-J.; Wang, T.; Chang, Z.-B.; Zeng, X.-Y. Key-Frame Extraction Based on HSV Histogram and Adaptive Clustering. Math. Probl. Eng. 2019, 2019, 1–10. [Google Scholar] [CrossRef]
- Zhao, B.; Xu, S.; Lin, S.; Wang, R.; Luo, X. A New Visual Interface for Searching and Navigating Slide-Based Lecture Videos. In Proceedings of the 2019 IEEE International Conference on Multimedia and Expo (ICME), Shanghai, China, 8–12 July 2019; IEEE: New York, NY, USA, 2019; pp. 928–933. [Google Scholar]
- Zhao, B.; Lin, S.; Luo, X.; Xu, S.; Wang, R. A Novel System for Visual Navigation of Educational Videos Using Multimodal Cues. In Proceedings of the 25th ACM International Conference on Multimedia, Mountain View, CA, USA, 23–27 October 2017; ACM: New York, NY, USA, 2017; pp. 1680–1688. [Google Scholar]
- Smeaton, A.F.; Over, P.; Doherty, A.R. Video Shot Boundary Detection: Seven Years of TRECVid Activity. Comput. Vis. Image Underst. 2010, 114, 411–418. [Google Scholar] [CrossRef]
- Wang, W.; Shen, J.; Shao, L. Consistent Video Saliency Using Local Gradient Flow Optimization and Global Refinement. IEEE Trans. Image Process. 2015, 24, 4185–4196. [Google Scholar] [CrossRef] [PubMed]
- Otani, M.; Nakashima, Y.; Rahtu, E.; Heikkilä, J.; Yokoya, N. Video Summarization Using Deep Semantic Features. In Proceedings of the 13th Asian Conference on Computer Vision, Taipei, Taiwan, 20–24 November 2016; Springer International Publishing: Cham, Switzerland, 2017; pp. 361–377. [Google Scholar]
- Repp, S.; Meinel, C. Semantic Indexing for Recorded Educational Lecture Videos. In Proceedings of the Fourth Annual IEEE International Conference on Pervasive Computing and Communications Workshops (PERCOMW’06), Pisa, Italy, 13–17 March 2006; IEEE: New York, NY, USA, 2006; pp. 240–245. [Google Scholar]
- Shiraiwa, T.; Nobuhara, H. Efficient Video Summarization Based on Semantic Segmentation Model. In Proceedings of the 2021 IEEE 10th Global Conference on Consumer Electronics (GCCE), Kyoto, Japan, 12–15 October 2021; IEEE: New York, NY, USA, 2021; pp. 436–439. [Google Scholar]
- Dhanushika, T.; Weerasinghe, T.A. Auto Identifying Key Information in Video Lectures to Generate a Navigation Structure. In Proceedings of the 2024 International Research Conference on Smart Computing and Systems Engineering (SCSE), Colombo, Sri Lanka, 4 April 2024; IEEE: New York, NY, USA, 2024; pp. 1–6. [Google Scholar]
- Dalal, N.; Triggs, B. Histograms of Oriented Gradients for Human Detection. In Proceedings of the 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition (CVPR’05), San Diego, CA, USA, 20–25 June 2005; IEEE: New York, NY, USA, 2005; pp. 886–893. [Google Scholar]
- Milosevic, N. Convolutions and Convolutional Neural Networks. In Introduction to Convolutional Neural Networks; Apress: Berkeley, CA, USA, 2020. [Google Scholar]
- De, P. Key Frame Extraction from Videos Based on SIFT and Structural Similarity. In International Conference on Communication and Intelligent System; Springer Nature: Singapore, 2023; pp. 361–371. [Google Scholar]
- Li, Y. Chitra Dorai SVM-Based Audio Classification for Instructional Video Analysis. In Proceedings of the 2004 IEEE International Conference on Acoustics, Speech, and Signal Processing, Montreal, QC, Canada, 17–21 May 2004; IEEE: New York, NY, USA, 2004; pp. V-897–900. [Google Scholar]
- Nandyal, S.; Kattimani, S.L. An Efficient Umpire Key Frame Segmentation in Cricket Video Using HOG and SVM. In Proceedings of the 2021 6th International Conference for Convergence in Technology (I2CT), Pune, India, 2–4 April 2021; IEEE: New York, NY, USA, 2021; pp. 1–7. [Google Scholar]
- Sun, Y.; Li, P.; Jiang, Z.; Hu, S. Feature Fusion and Clustering for Key Frame Extraction. Math. Biosci. Eng. 2021, 18, 9294–9311. [Google Scholar] [CrossRef] [PubMed]
- Hashemi, N.S.; Aghdam, R.B.; Ghiasi, A.S.B.; Fatemi, P. Template Matching Advances and Applications in Image Analysis. arXiv 2016, arXiv:1610.07231. [Google Scholar] [CrossRef]
- Castillo-Abdul, B.; González-Carrión, E.L.; Cabrero, J.D.B. Involvement of Moocs in The Teachinglearning Process. J. Entrep. Educ. 2021, 24, 1–10. [Google Scholar]
- Zhang, X.; Li, C.; Li, S.-W.; Zue, V. Automated Segmentation of MOOC Lectures towards Customized Learning. In Proceedings of the 2016 IEEE 16th International Conference on Advanced Learning Technologies (ICALT), Austin, TX, USA, 25–28 July 2016; IEEE: New York, NY, USA 2016; pp. 20–22. [Google Scholar]
- Baidya, E.; Goel, S. LectureKhoj: Automatic Tagging and Semantic Segmentation of Online Lecture Videos. In Proceedings of the 2014 Seventh International Conference on Contemporary Computing (IC3), Noida, India, 7–9 August 2014; IEEE: New York, NY, USA, 2014; pp. 37–43. [Google Scholar]
- Shah, R.R.; Yu, Y.; Shaikh, A.D.; Tang, S.; Zimmermann, R. ATLAS. In Proceedings of the 22nd ACM international conference on Multimedia, Orlando, FL, USA, 3–7 November 2017; ACM: New York, NY, USA, 2014; pp. 209–212. [Google Scholar]
- Biswas, A.; Gandhi, A.; Deshmukh, O. MMToC. In Proceedings of the 23rd ACM International Conference on Multimedia, Brisbane, QLD, Australia, 13–17 October 2015; ACM: New York, NY, USA, 2015; pp. 621–630. [Google Scholar]
- Mahapatra, D.; Mariappan, R.; Rajan, V. Automatic Hierarchical Table of Contents Generation for Educational Videos. In Proceedings of the Companion of the Web Conference 2018—WWW ’18, Lyon, France, 23–27 April 2018; ACM Press: New York, NY, USA, 2018; pp. 267–274. [Google Scholar]
- Mahapatra, D.; Mariappan, R.; Rajan, V.; Yadav, K.A.; Roy, S. VideoKen. In Proceedings of the Companion of the Web Conference 2018—WWW ’18, Lyon, France, 23–27 April 2018; ACM Press: New York, NY, USA, 2018; pp. 239–242. [Google Scholar]
- van den Oord, A.; Li, Y.; Vinyals, O. Representation Learning with Contrastive Predictive Coding. arXiv 2019, arXiv:1807.03748. [Google Scholar] [CrossRef]
- Soares, E.R.; Barrére, E. An Optimization Model for Temporal Video Lecture Segmentation Using Word2vec and Acoustic Features. In Proceedings of the 25th Brazillian Symposium on Multimedia and the Web, Rio de Janeiro, Brazil, 29 October–1 November 2019; ACM: New York, NY, USA, 2019; pp. 513–520. [Google Scholar]
- Liu, T.; Choudary, C. Content Extraction and Summarization of Instructional Videos. In Proceedings of the 2006 International Conference on Image Processing, Atlanta, GA, USA, 8–11 October 2006; IEEE: Piscataway, NJ, USA, 2006; pp. 149–152. [Google Scholar]
- Wang, Z.; Zhang, M.; Chen, H.; Li, J.; Li, G.; Zhao, J.; Yao, L.; Zhang, J.; Chu, F. A Generalized Fault Diagnosis Framework for Rotating Machinery Based on Phase Entropy. Reliab. Eng. Syst. Saf. 2025, 256, 110745. [Google Scholar] [CrossRef]
- Younes, A.; Schaub-Meyer, S.; Chalvatzaki, G. Entropy-Driven Unsupervised Keypoint Representation Learning in Videos. In Proceedings of the 2023 International Conference on Machine Learning, Honolulu, HI, USA, 23–29 July 2023; PMLR: New York, NY, USA, 2023. [Google Scholar]
- Su, R.; Huang, W.; Ma, H.; Song, X.; Hu, J. SGE Net: Video Object Detection with Squeezed GRU and Information Entropy Map. In Proceedings of the 2021 IEEE International Conference on Image Processing (ICIP), Anchorage, AK, USA, 19–22 September 2021. [Google Scholar]
- Zhang, X.; Fu, D.; Liu, N. Shot Segmentation Based on Von Neumann Entropy for Key Frame Extraction. arXiv 2024, arXiv:2408.15844. [Google Scholar] [CrossRef]
- Masry, A.; Long, D.X.; Tan, J.Q.; Joty, S.; Hoque, E. ChartQA: A Benchmark for Question Answering about Charts with Visual and Logical Reasoning. arXiv 2022, arXiv:2203.10244. [Google Scholar] [CrossRef]
- Kembhavi, A.; Salvato, M.; Kolve, E.; Seo, M.; Hajishirzi, H.; Farhadi, A. A Diagram Is Worth a Dozen Images. In Proceedings of the 2016 European Conference on Computer Vision, Amsterdam, The Netherlands, 8–16 October 2016; Springer International Publishing: Cham, Switzerland, 2016; pp. 235–251. [Google Scholar]
- Piasco, N.; Sidibé, D.; Gouet-Brunet, V.; Demonceaux, C. Improving Image Description with Auxiliary Modality for Visual Localization in Challenging Conditions. Int. J. Comput. Vis. 2021, 129, 185–202. [Google Scholar] [CrossRef]
- Smith, R. An Overview of the Tesseract OCR Engine. In Proceedings of the Ninth International Conference on Document Analysis and Recognition (ICDAR 2007), Paraná, Brazil, 23–26 September 2007; IEEE: New York, NY, USA, 2007; Volume 2, pp. 629–633. [Google Scholar]
- Zellers, R.; Lu, J.; Lu, X.; Yu, Y.; Zhao, Y.; Salehi, M.; Kusupati, A.; Hessel, J.; Farhadi, A.; Choi, Y. MERLOT RESERVE: Neural Script Knowledge through Vision and Language and Sound. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 19–24 June 2022; IEEE: New York, NY, USA 2022; pp. 16354–16366. [Google Scholar]
- Li, Z.; Gavrilyuk, K.; Gavves, E.; Jain, M.; Snoek, C.G. The How2 Dataset and Multimodal Baselines. In Proceedings of the 2020 International Conference on Multimedia Retrieval, Dublin, Ireland, 26–29 October 2020; pp. 287–294. [Google Scholar]
- Grauman, K.; Westbury, A.; Byrne, E.; Chavis, Z.; Furnari, A.; Girdhar, R.; Hamburger, J.; Jiang, H.; Liu, M.; Liu, X.; et al. Ego4d: Around the World in 3000 h of Egocentric Video. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, New Orleans, LA, USA, 19–24 June 2022; pp. 18995–19012. [Google Scholar]
- NPTEL Online Certification. Indian Institute of Technology & Indian Institute of Science. National Programme on Technology Enhanced Learning (NPTEL). Available online: https://nptel.ac.in/ (accessed on 9 October 2025).
- Haurilet, M.; Al-Halah, Z.; Stiefelhagen, R. SPaSe–Multi-Label Page Segmentation for Presentation Slides. In Proceedings of the 2019 IEEE Winter Conference on Applications of Computer Vision (WACV), Waikoloa Village, HI, USA, 7–11 January 2019; IEEE: New York, NY, USA, 2019; pp. 726–734. [Google Scholar]
- Haurilet, M.; Roitberg, A.; Martinez, M.; Stiefelhagen, R. WiSe—Slide Segmentation in the Wild. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR), Sydney, NSW, Australia, 20–25 September 2019; IEEE: New York, NY, USA, 2019; pp. 343–348. [Google Scholar]
- Dutta, K.; Mathew, M.; Krishnan, P.; Jawahar, C.V. Localizing and Recognizing Text in Lecture Videos. In Proceedings of the 2018 16th International Conference on Frontiers in Handwriting Recognition (ICFHR), Niagara Falls, NY, USA, 5–8 August 2018; IEEE: New York, NY, USA, 2018; pp. 235–240. [Google Scholar]
- Wang, W.; Song, Y.; Jha, S. Autolv: Automatic Lecture Video Generator. In Proceedings of the 2022 IEEE International Conference on Image Processing (ICIP), Bordeaux, France, 16–19 October 2022; IEEE: New York, NY, USA, 2022; pp. 1086–1090. [Google Scholar]
- Lee, D.W.; Ahuja, C.; Liang, P.P.; Natu, S.; Morency, L.-P. Multimodal Lecture Presentations Dataset: Understanding Multimodality in Educational Slides. arXiv 2022, arXiv:2208.08080. [Google Scholar] [CrossRef]
- Li, I.; Fabbri, A.R.; Tung, R.R.; Radev, D.R. What Should I Learn First: Introducing LectureBank for NLP Education and Prerequisite Chain Learning. Proc. AAAI Conf. Artif. Intell. 2019, 33, 6674–6681. [Google Scholar] [CrossRef]
- Bulathwela, S.; Perez-Ortiz, M.; Yilmaz, E.; Shawe-Taylor, J. VLEngagement: A Dataset of Scientific Video Lectures for Evaluating Population-Based Engagement. arXiv 2020, arXiv:2011.02273. [Google Scholar] [CrossRef]
- Araujo, A.Y.M.; G.B. ClassX Dataset (Classx, Video Search, Query-by-Image, and Lecture Videos). Available online: https://Exhibits.Stanford.Edu/Data/Catalog/Sf888mq5505 (accessed on 13 December 2023).
- Minderer, M.; Gritsenko, A.; Stone, A.; Neumann, M.; Weissenborn, D.; Dosovitskiy, A.; Mahendran, A.; Arnab, A.; Dehghani, M.; Shen, Z.; et al. Simple Open-Vocabulary Object Detection with Vision Transformers. In Proceedings of the 2022 European Conference on Computer Vision, Tel Aviv, Israel, 23–27 October 2022; Springer Nature: Cham, Switzerland, 2022. [Google Scholar]
- Heigold, G.; Keysers, D.; Minderer, M.; Lučić, M.; Gritsenko, A.; Yu, F.; Bewley, A.; Kipf, T. Video OWL-ViT: Temporally-Consistent Open-World Localization in Video. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; IEEE: New York, NY, USA, 2023; pp. 13756–13765. [Google Scholar]
- Li, L.H.; Zhang, P.; Zhang, H.; Yang, J.; Li, C.; Zhong, Y.; Wang, L.; Yuan, L.; Zhang, L.; Hwang, J.-N.; et al. Grounded Language-Image Pre-Training. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; IEEE: New York, NY, USA, 2022; pp. 10955–10965. [Google Scholar]
- Kirillov, A.; Mintun, E.; Ravi, N.; Mao, H.; Rolland, C.; Gustafson, L.; Xiao, T.; Whitehead, S.; Berg, A.C.; Lo, W.-Y.; et al. Segment Anything. In Proceedings of the 2023 IEEE/CVF International Conference on Computer Vision (ICCV), Paris, France, 1–6 October 2023; IEEE: New York, NY, USA, 2023; pp. 3992–4003. [Google Scholar]
- Ren, T.; Jiang, Q.; Liu, S.; Zeng, Z.; Liu, W.; Gao, H.; Huang, H.; Ma, Z.; Jiang, X.; Chen, Y.; et al. Grounding DINO 1.5: Advance the “Edge” of Open-Set Object Detection. arXiv 2024, arXiv:2405.10300. [Google Scholar] [CrossRef]
- Zhu, X.; Su, W.; Lu, L.; Li, B.; Wang, X.; Dai, J. Deformable DETR: Deformable Transformers for End-to-End Object Detection. arXiv 2020, arXiv:2010.04159. [Google Scholar]
- Mal, Z.; Luo, G.; Gao, J.; Li, L.; Chen, Y.; Wang, S.; Zhang, C.; Hu, W. Open-Vocabulary One-Stage Detection with Hierarchical Visual-Language Knowledge Distillation. In Proceedings of the 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), New Orleans, LA, USA, 18–24 June 2022; IEEE: New York, NY, USA, 2022; pp. 14054–14063. [Google Scholar]
- Li, M.; Cui, L.; Huang, S.; Wei, F.; Zhou, M.; Li, Z. TableBank: Table Benchmark for Image-Based Table Detection and Recognition. In Proceedings of the Twelfth Language Resources and Evaluation Conference, Marseille, France, 11–16 May 2020; Calzolari, N., Béchet, F., Blache, P., Choukri, K., Cieri, C., Declerck, T., Goggi, S., Isahara, H., Maegaard, B., Mariani, J., et al., Eds.; European Language Resources Association: Luxembourg, 2020; pp. 1918–1925. [Google Scholar]
- Haloi, M.; Shekhar, S.; Fande, N.; Dash, S.S.; G, S. Table Detection in the Wild: A Novel Diverse Table Detection Dataset and Method. arXiv 2023, arXiv:2209.09207. [Google Scholar] [CrossRef]
- Hands, A. Duckduckgo http://www.duckduckgo.com or http://www.ddg.gg. Tech. Serv. Q. 2012, 29, 345–347. [Google Scholar] [CrossRef]
- Araujo, A.; Chaves, J.; Lakshman, H.; Angst, R.; Girod, B. Large-Scale Query-by-Image Video Retrieval Using Bloom Filters. arXiv 2016, arXiv:1604.07939. [Google Scholar]
- Paliwal, S.S.; D, V.; Rahul, R.; Sharma, M.; Vig, L. TableNet: Deep Learning Model for End-to-End Table Detection and Tabular Data Extraction from Scanned Document Images. In Proceedings of the 2019 International Conference on Document Analysis and Recognition (ICDAR), Sydney, NSW, Australia, 20–25 September 2019; IEEE: New York, NY, USA, 2019; pp. 128–133. [Google Scholar]
- Roboflow: End-to-End Computer Vision Platform. Available online: https://universe.roboflow.com/ (accessed on 9 October 2025).
- Lei, J.; Yu, L.; Bansal, M.; Berg, T.L. TVQA: Localized, Compositional Video Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium, 31 October–4 November 2019. [Google Scholar]
- Miech, A.; Zhukov, D.; Alayrac, J.-B.; Tapaswi, M.; Laptev, I.; Sivic, J. HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips. In Proceedings of the 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Republic of Korea, 29 October–1 November 2019; IEEE: New York, NY, USA, 2019; pp. 2630–2640. [Google Scholar]
- Tang, Y.; Ding, D.; Rao, Y.; Zheng, Y.; Zhang, D.; Zhao, L.; Lu, J.; Zhou, J. COIN: A Large-Scale Dataset for Comprehensive Instructional Video Analysis. In Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, 15–20 June 2019; IEEE: New York, NY, USA, 2019; pp. 1207–1216. [Google Scholar]
- Davila, K.; Xu, F.; Setlur, S.; Govindaraju, V. FCN-LectureNet: Extractive Summarization of Whiteboard and Chalkboard Lecture Videos. IEEE Access 2021, 9, 104469–104484. [Google Scholar] [CrossRef]
- Sharma, V.; Gupta, M.; Kumar, A.; Mishra, D. EduNet: A New Video Dataset for Understanding Human Activity in the Classroom Environment. Sensors 2021, 21, 5699. [Google Scholar] [CrossRef]
- Lin, T.-Y.; Goyal, P.; Girshick, R.; He, K.; Dollár, P. Focal Loss for Dense Object Detection. In Proceedings of the IEEE International Conference on Computer Vision, Venice, Italy, 22–29 October 2018. [Google Scholar]
- de Avila, S.E.F.; Lopes, A.P.B.; da Luz, A.; de Albuquerque Araújo, A. VSUMM: A Mechanism Designed to Produce Static Video Summaries and a Novel Evaluation Method. Pattern Recognit. Lett. 2011, 32, 56–68. [Google Scholar] [CrossRef]
- Ren, S.; He, K.; Girshick, R.; Sun, J. Faster R-CNN: Towards Real-Time Object Detection with Region Proposal Networks. IEEE Trans. Pattern Anal. Mach. Intell. 2017, 39, 1137–1149. [Google Scholar] [CrossRef]
- Biswas, D.; Shah, S.; Subhlok, J. Lecture Video Visual Objects (LVVO) Dataset: A Benchmark for Visual Object Detection in Educational Videos. arXiv 2025, arXiv:2506.13657. [Google Scholar] [CrossRef]
- Xue, H.; Sun, Y.; Liu, B.; Fu, J.; Song, R.; Li, H.; Luo, J. CLIP-ViP: Adapting Pre-Trained Image-Text Model to Video-Language Representation Alignment. arXiv 2023, arXiv:2209.06430. [Google Scholar]
- OpenAI; Achiam, J.; Adler, S.; Agarwal, S.; Ahmad, L.; Akkaya, I.; Aleman, F.L.; Almeida, D.; Altenschmidt, J.; Altman, S.; et al. GPT-4 Technical Report. arXiv 2024, arXiv:2303.08774. [Google Scholar]
- Xu, H.; Ghosh, G.; Huang, P.-Y.; Okhonko, D.; Aghajanyan, A.; Metze, F.; Zettlemoyer, L.; Feichtenhofer, C. VideoCLIP: Contrastive Pre-Training for Zero-Shot Video-Text Understanding. arXiv 2021, arXiv:2109.14084. [Google Scholar]
- Adcock, J.; Cooper, M.; Denoue, L.; Pirsiavash, H.; Rowe, L.A. TalkMiner. In Proceedings of the Proceedings of the 18th ACM International Conference on Multimedia, Florence, Italy, 25–29 October 2010; ACM: New York, NY, USA, 2010; pp. 241–250. [Google Scholar]
- Dhanushree, M.; Priya, R.; Aruna, P.; Bhavani, R. A Framework for Video Summarization Using Visual Attention Technique. Indian J. Sci. Technol. 2024, 17, 1586–1595. [Google Scholar] [CrossRef]
- Biswas, D.; Shah, S.; Subhlok, J. Visual Content Detection in Educational Videos with Transfer Learning and Dataset Enrichment. arXiv 2025, arXiv:2506.21903. [Google Scholar] [CrossRef]
- Psallidas, T.; Spyrou, E. Video Summarization Based on Feature Fusion and Data Augmentation. Computers 2023, 12, 186. [Google Scholar] [CrossRef]



















| Dataset Name | Slide Segment | Slide Figures | Slide Text | Spoken Language | Relationship Annotation | #Videos | #Hours | #Slides | Availability |
|---|---|---|---|---|---|---|---|---|---|
| ClassX [50] | ✓ | - | - | - | - | - | 408 | 1,478,978 | ✓ |
| NPTEL [42] | - | - | - | ✓ (transcribed) | - | 10,000+ | 56,000+ | - | ✓ |
| SpaSe [43] | ✓ | - | - | - | - | - | - | 2000 | ✓ |
| WiSe [44] | ✓ | ✓ | ✓ | - | - | - | - | 1300 | ✓ |
| Lecture VideoDB [45] | ✓ | - | ✓ | - | - | 24 | - | 5000 | ✓ |
| ALV [46] | ✓ | - | - | ✓ | - | - | - | 1498 | ✓ |
| MLP [47] | ✓ | ✓ | ✓ | ✓ | - | 334 | 187 | 9031 | ✓ |
| LectureBank [48] | ✓ | - | ✓ | - | ✓(prerequisite) | 1352 | - | 51,939 | ✓ |
| Educational Stage | Primary Content Type | Knowledge Density and Presentation Style | Representative Examples from EVUD-2M | Model Applicability and Notes |
|---|---|---|---|---|
| Undergraduate and Postgraduate University Courses | Theoretical and formal instruction (primary focus) | High knowledge density. Formal presentation of concepts using slides, black/whiteboards. Characterized by mathematical formulations, complex diagrams, and structured arguments. |
| Strong applicability. LEARNet is explicitly designed for this “Instructional Steady-State” with static, information-rich visuals. |
| Advanced Professional Development | Specialized technical training | Moderate to High knowledge density. Similarly to postgraduate content, often using slide decks. Focuses on advanced, industry-relevant topics. |
| Strong applicability. Shares the visual and structural characteristics of university-level lectures. |
| Vocational/Practical Demonstration | Physical skill demonstration (currently under-represented) | Low to Moderate conceptual density, High action density. Information is conveyed through physical actions, tool use, and dynamic processes. Minimal reliance on static text/diagrams. | Limited examples (e.g., general “Aircraft Maintenance”). Lacks specific, action-centric videos like “Chemistry Experiments” or “Mechanical Operation”. | Limited current applicability. A key area for future dataset expansion to handle dynamic, non-lecture formats. |
| Dataset | Keyframe Extraction | Region Annotations | Hybrid Content Support | One-Shot Capability | Low-Motion Optimization |
|---|---|---|---|---|---|
| TVQA [64] | ✖ | ✖ | ✖ (TV episodes) | ✖ | ✖ |
| HowTo100M [65] | ✔ (aligned with speech) | ✖ | ✖ | ✖ | ✖ |
| COIN [66] | ✔ (action steps) | ✖ | ✖ (instructional videos) | ✖ | ✖ |
| LectureNet [67] | ✔ | ✖ | Partially | ✖ | ✖ |
| EduNet [68] | ✔ | ✖ | Partially | ✖ | ✖ |
| ChartQA [35]/AI2D [36] | ✖ | ✔ (diagrams) | ✖ | ✖ | ✖ |
| EVUD-2M (Ours) | ✔ (semantically filtered) | ✔ (educational) | ✔ | ✔ | ✔ (designed for lecture video) |
| Component | Setting |
|---|---|
| Input region size | 224 × 224 × 3 |
| Backbone conv layers | 3 × Conv (3 × 3, BN, ReLU) |
| Conv filter sizes | 64 → 128 → 256 |
| Feature aggregation | Global Average Pooling (512D descriptor) |
| Projection layer | 512 → 512-D |
| Embedding normalization | L2 normalization |
| Similarity metric | Cosine similarity |
| Decision rule | Similarity threshold () |
| Reference set size | 200 exemplar regions |
| Loss function | Triplet Loss |
| Margin (m) | 0.2 |
| Optimizer | AdamW |
| Learning rate | 1 × 10−3 |
| Weight decay | 0.05 |
| Batch size | 32 region triples |
| Number of epochs | 50 |
| Sampling strategy | Hard Negative Mining |
| Training objective | Relational consistency embedding learning |
| Weight α | F1-Score (Keyframe Selection) | Interpretation |
|---|---|---|
| 0.67 | 0.871 (Peak) | Achieves the best balance between structural change (67%) and saliency, yielding the most pedagogically meaningful keyframe selection. |
| 0.5 | 0.812 | Underweights structural cues, resulting in redundant and visually similar keyframes. |
| 0.6 | 0.855 | Provides moderate balance but does not reach the optimal trade-off achieved at α = 0.67. |
| 0.7 | 0.86 | Slightly overemphasizes structural transitions, leading to minor loss of salient instructional content. |
| 0.8 | 0.84 | Overweight’s structural change and suppresses important salient regions, decreasing overall accuracy. |
| PIS Weight (α) | Extracted Keyframes | DPC Cut-Off Parameter (σ) | Observation |
|---|---|---|---|
| 0.50 | 185 | 4.9509 | Compressed PIS → smallest σ |
| 0.60 | 180 | 5.7572 | Slightly wider PIS → moderate σ |
| 0.67 | 181 | 6.3596 | Balanced PIS → moderate σ (optimal) |
| 0.70 | 182 | 6.6241 | Slightly wider → higher σ |
| 0.80 | 190 | 7.5303 | Widest PIS → largest σ, mild smoothing |
| Method | Precision (%) | Recall (%) | F1-Score | (%) | Limitations |
|---|---|---|---|---|---|
| VSUMM [70] | 83.7 | 80.3 | 0.82 | 52.1 | Misses content-rich frames without motion |
| Motion-based [17] | 83.0 | 79.0 | 0.81 | ≤50 | Ignores semantic relevance; unsuitable for static content |
| CLIP + Optical flow [21] | 86.8 | 85.0 | 0.86 | 65.5 | Computationally heavy (3.2 sec/frame); sensitive to text density |
| TIB (Ours) | 89.0 | 88.0 | 0.89 | 70.2 | Designed for pedagogical transitions |
| Filtering Stage | Frames Discarded | Educational Rationale |
|---|---|---|
| Blank/Logo Detection | 42% | Eliminates non-instructional content |
| Talking-Head Removal | 31% | Focuses on educational visual aids |
| High-Similarity Frames | 27% | Reduces redundant content |
| Final Reduction | 70.2% | Overall pedagogical efficiency |
| Method | Approach | Expected Accuracy/Performance | Notes |
|---|---|---|---|
| TalkMiner [76] | Slide detection + OCR; selects slide keyframes and indexes text | 0.76 | Effective for slide-heavy lectures; limited for whiteboard content, handwritten notes, and dynamic diagrams; dependent on OCR quality. |
| Manual Annotation Screening | Human experts select pedagogically relevant keyframes | 0.89 | High accuracy but not scalable; labor-intensive and impractical for large datasets. |
| Keyword Matching/TF-IDF [77] | Extracts frames containing specific textual keywords from slides or transcripts | 0.62 | Misses diagrams, formulas, and non-text visuals; biased toward text-rich slides. |
| Saliency-Based Selection [78] | Uses visual saliency maps to identify visually “distinctive” frames | 0.68 | Captures high-contrast regions but ignores pedagogical relevance; misses subtle instructional transitions. |
| Entropy Reduction (Proposed) | Automatically selects frames with high visual information (slides, diagrams, tables, equations) using entropy metrics | 0.87 (Keyframe F1-Score, measured on EVUD-2M) | Captures diagrams, tables, equations, and mixed-format visuals; scalable and domain-aware across diverse lecture styles. |
| Method | mAP@50 | Generalization to Unseen Symbols | Core Educational Limitation |
|---|---|---|---|
| Faster R-CNN [71] | 0.66 | ✗ | Restricted vocabulary; unable to detect new elements |
| Grounding DINO [55] | 0.68 | ✓ | Inconsistent precision for visually dense instructional regions |
| DETR- 50 [56] | 0.62 | ✗ | Struggles with fine-grained educational structures |
| OWL-ViT [51] | 0.71 | ✓ | Coarse spatial precision for educational regions |
| LEARNet (Ours) | 0.88 | ✓ | High-precision parsing of educational regions |
| Model | mAP@50 (LVVO Test Set) | Output Type | Keyframe Selection | Relational Consistency Check |
|---|---|---|---|---|
| LVVO (Faster R-CNN) | 0.72 | Bounding Boxes | ✗ | ✗ |
| LEARNet (SSD Module) | 0.86 | Pixel Masks | ✓ | ✗ |
| LEARNet (Full Framework) | 0.88 * | Masks + RCVN | ✓ | ✓ |
| Media Type | Precision | Recall | F1-Score | Educational Challenges Addressed |
|---|---|---|---|---|
| Slides | 0.90 | 0.89 | 0.895 | Structured content preservation |
| Blackboard | 0.89 | 0.87 | 0.879 | Handwriting variability |
| Digital Writing | 0.79 | 0.67 | 0.725 | Dynamic stroke complexity |
| Category | Precision | Recall | F1-Score | mAP@50 | Key Challenges |
|---|---|---|---|---|---|
| Text_Block | 0.81 | 0.78 | 0.79 | 0.79 | Dense layout and small font size affect recall in packed slides. |
| Handwritten Text | 0.74 | 0.68 | 0.71 | 0.71 | High stroke variability, noise, and low contrast on digital whiteboards. |
| Diagram/ Flowchart | 0.90 | 0.88 | 0.89 | 0.89 | Clear visual boundaries and distinct graphical nature aid detection. |
| Formula/Equation | 0.82 | 0.79 | 0.80 | 0.81 | Complex symbol relationships and occasional inconsistent formatting. |
| Table | 0.93 | 0.91 | 0.92 | 0.92 | Regular, explicit structural cues enable high accuracy. |
| Slide_Title | 0.88 | 0.85 | 0.86 | 0.86 | High contrast and large font size contribute to reliable detection. |
| Code_Block | 0.85 | 0.82 | 0.83 | 0.83 | Monospaced font and box delineation provide strong visual cues. |
| Figure | 0.78 | 0.72 | 0.75 | 0.75 | Lower performance due to high internal visual variation (generic images). |
| Component | mAP@50 | False Positives | Educational Impact |
|---|---|---|---|
| Replacing TIB with FFmpeg keyframe extraction | 0.68 | +62% | Retains many irrelevant frames; misses subtle instructional transitions. |
| With TIB (Our Temporal Module) | 0.88 | Baseline | Highly selective; filters redundant/non-pedagogical frames. |
| Replacing SSD with OWL-ViT | 0.71 | +48% | Coarse localization increases mis-detections, especially for thin or irregular elements. |
| Removing RCVN | 0.79 | +32% | Semantic inconsistencies rise; relational errors cause more mismatched labels. |
| Full LEARNet (TIB + SSD + RCVN) | 0.88 | Baseline | Optimal temporal, spatial, and semantic precision with minimal erroneous detections. |
| Model | mAP@50 | Computational Profile | Fundamental Educational Misalignment |
|---|---|---|---|
| GPT-4V [74] | 0.72 | High inference latency (~2.1s/frame); imprecise bounding boxes; poor spatial accuracy | Ignores Spatial and Temporal Symmetry: Processes frames in isolation, blind to consistent slide layouts and repetitive element arrangements. |
| CLIP-ViP [73] | 0.75 | Sensitive to prompt phrasing; fails to distinguish semantically different but visually similar elements | Breaks Semantic Symmetry: Relies on superficial visual-textual correlations, missing pedagogically significant subtle changes |
| VideoCLIP [75] | 0.71 | Optimized for dynamic action; poor performance on static, slide-based content | Violates Temporal Symmetry: Action-oriented temporal modeling misaligned with structured lecture pacing |
| LEARNet (Ours) | 0.88 | Superior performance and efficiency; education-domain optimized | Explicitly Models Symmetry: TIB preserves temporal–semantic coherence; SSD+RCVN enforce spatial and relational consistency. |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2025 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
S, C.; V V, N.; S R, N. LEARNet: A Learning Entropy-Aware Representation Network for Educational Video Understanding. Entropy 2026, 28, 3. https://doi.org/10.3390/e28010003
S C, V V N, S R N. LEARNet: A Learning Entropy-Aware Representation Network for Educational Video Understanding. Entropy. 2026; 28(1):3. https://doi.org/10.3390/e28010003
Chicago/Turabian StyleS, Chitrakala, Nivedha V V, and Niranjana S R. 2026. "LEARNet: A Learning Entropy-Aware Representation Network for Educational Video Understanding" Entropy 28, no. 1: 3. https://doi.org/10.3390/e28010003
APA StyleS, C., V V, N., & S R, N. (2026). LEARNet: A Learning Entropy-Aware Representation Network for Educational Video Understanding. Entropy, 28(1), 3. https://doi.org/10.3390/e28010003

