Next Article in Journal
Volt–Var Self-Optimizing Control of Distribution Networks Based on the BOST-GRPO Algorithm Under Stability Constraints
Previous Article in Journal
A Real-World Benchmark for Early Wildfire Detection Using Sequential Data with the PyroNear Dataset
Previous Article in Special Issue
Time-Series Modeling Based on a Modified Volterra Neural Network
 
 
Article
Peer-Review Record

Convolutional Neural Networks: Biological Foundations, Hidden Limitations, and Future Directions

Electronics 2026, 15(12), 2654; https://doi.org/10.3390/electronics15122654
by Luis Sacouto 1,* and Andreas Wichert 2
Reviewer 1:
Reviewer 2: Anonymous
Reviewer 3:
Electronics 2026, 15(12), 2654; https://doi.org/10.3390/electronics15122654
Submission received: 14 May 2026 / Revised: 4 June 2026 / Accepted: 8 June 2026 / Published: 15 June 2026
(This article belongs to the Special Issue Convolutional Neural Networks and Vision Applications, 4th Edition)

Round 1

Reviewer 1 Report

Comments and Suggestions for Authors
  1. The manuscript presents an interesting and intellectually provocative perspective on CNN architectures by revisiting their biological foundations and connecting them to known geometric limitations. The paper is well organized and technically ambitious, especially in its attempt to unify neuroscience, computational vision, and deep learning theory into a coherent critique.

  2. Although the manuscript is conceptually strong, the writing is excessively verbose throughout multiple sections. Many arguments are repeated with slightly different wording, particularly regarding pooling, geometry loss, and modality conflation. The paper would benefit significantly from condensation and tighter presentation.

  3. The central thesis regarding the “pooling-as-complex-cell category error” is interesting, but several claims are presented too strongly without sufficient experimental validation. The manuscript relies heavily on conceptual and theoretical reasoning while offering limited empirical evidence directly supporting the proposed conclusions.

  4. The discussion about the biological hierarchy and the claim that CNNs incorrectly extend the S-C operation beyond V3 is thought-provoking. However, the manuscript occasionally overstates biological certainty while simplifying the complexity of higher visual cortex processing. Additional caution and nuance are needed when drawing one-to-one architectural comparisons between biological and artificial systems.

  5. Figure 1 effectively summarizes the invariance-discrimination tradeoff and provides a strong conceptual anchor for the paper. Similarly, Figures 2 and 3 are clear and helpful in illustrating the distinction between geometry-preserving operations and pooling-based information loss.

  6. The paper critiques CNNs extensively, but the discussion of modern architectures such as ConvNeXt, hybrid ViT-CNN models, Mamba-based vision architectures, hierarchical transformers, and diffusion-based representations is relatively limited. Including these developments would improve the completeness and timeliness of the review.

  7. The manuscript frequently frames pooling as fundamentally destructive to geometry. While this argument is theoretically motivated, the practical success of deep CNNs in dense prediction tasks such as segmentation, detection, and medical imaging suggests the issue may be more nuanced. The discussion should acknowledge this more explicitly.

  8. The section discussing adversarial vulnerability is interesting, but the connection between geometric representation and adversarial robustness remains largely speculative. More direct evidence or references supporting this causal relationship would strengthen the argument.

  9. The review would benefit from a dedicated section discussing whether current architectural remedies such as skip connections, feature pyramids, attention mechanisms, multi-scale fusion, or positional encoding partially mitigate the identified geometric degradation problem.

  10. Some arguments against pooling appear to overlook the fact that many modern architectures already reduce or eliminate aggressive pooling operations in favor of strided convolutions, adaptive feature aggregation, or patch-based tokenization. This evolution in the field should be discussed more comprehensively.

  11. The manuscript is academically rich, but several paragraphs contain very long sentences that reduce readability. Breaking these into shorter statements would improve clarity and accessibility for a broader readership.

  12. The proposed six desiderata are conceptually meaningful, but they remain relatively abstract. The paper would be strengthened by providing a more concrete architectural framework or preliminary design strategy demonstrating how these principles could realistically be implemented in practice.

  13. DOI: 10.1109/TGRS.2026.3685508 could strengthen the discussion regarding geometry-aware representation learning and structural feature preservation in deep architectures. Additionally, https://doi.org/10.1007/s13244-018-0639-9 may support the discussion on hybrid deep learning frameworks and limitations of conventional CNN feature representations.

Author Response

### Response to Reviewer 1
We thank the reviewer for the careful and constructive engagement with the manuscript. We note with appreciation the positive assessments in Comments 1 and 5, and respond to each critical comment below.
---
**Comments 1:** The manuscript presents an interesting and intellectually provocative perspective on CNN architectures by revisiting their biological foundations and connecting them to known geometric limitations. The paper is well organized and technically ambitious, especially in its attempt to unify neuroscience, computational vision, and deep learning theory into a coherent critique.
**Response 1:** We thank the reviewer for this assessment. No changes were made in response to this comment.
---
**Comments 2:** Although the manuscript is conceptually strong, the writing is excessively verbose throughout multiple sections. Many arguments are repeated with slightly different wording, particularly regarding pooling, geometry loss, and modality conflation. The paper would benefit significantly from condensation and tighter presentation.
**Response 2:** We have reviewed the manuscript for genuinely redundant passages and tightened several of them. One instance of the "coherent package" phrase, which appeared twice in close proximity in Section 9 (page 22), has been removed from the first occurrence. We note, however, that some recurrence of the core argument is structural rather than incidental: the same architectural failure, namely the pooling-as-complex-cell misconception, is examined successively from a biological perspective (Section 2), a computational perspective (Section 5), an empirical-consequences perspective (Section 6), and a design-implications perspective (Section 9). Each context develops the argument in ways that are not present in the others, and collapsing these appearances would sacrifice the progression from diagnosis to implication that motivates the paper's structure.
---
**Comments 3:** The central thesis regarding the "pooling-as-complex-cell category error" is interesting, but several claims are presented too strongly without sufficient experimental validation. The manuscript relies heavily on conceptual and theoretical reasoning while offering limited empirical evidence directly supporting the proposed conclusions.
**Response 3:** The central thesis is a conceptual analysis grounded in a formal result (Barnard & Casasent 1990) and in the established literature on complex cell physiology (Adelson & Bergen 1985). The empirical claims that accompany the thesis, specifically texture bias, out-of-distribution fragility, and adversarial vulnerability, are not presented as novel findings but as documented phenomena whose structural origin the paper diagnoses. They are supported by extensive cited empirical literature, including Geirhos et al. (2018, 2019), Baker et al. (2018), Hermann et al. (2019), Arjovsky (2021), and Krueger et al. (2021). The paper is explicit about the boundaries of what it claims: where the causal chain has not been empirically established at scale, the manuscript says so. This is consistent with the nature of a literature review article whose contribution is synthesis and diagnosis rather than new measurement. No manuscript changes were made in response to this comment.
---
**Comments 4:** The discussion about the biological hierarchy and the claim that CNNs incorrectly extend the S-C operation beyond V3 is thought-provoking. However, the manuscript occasionally overstates biological certainty while simplifying the complexity of higher visual cortex processing. Additional caution and nuance are needed when drawing one-to-one architectural comparisons between biological and artificial systems.
**Response 4:** We agree that additional hedging was warranted on the V3 boundary claim. In response to the same comment from Reviewer 3, the phrase "stops applying altogether" beyond V3 was overstated and has been revised to "becomes progressively less adequate as a characterization of the computational principles employed." A primary source (DiCarlo, Zoccolan & Rust, 2012, *Neuron* 73:415–434) has been added for the V4/IT selectivity claim. This change appears in Section 2 (page 4), in the paragraph discussing the S-C organizational principle. The revised passage reads:
> "…beyond V3 the Hubel-Wiesel framework becomes progressively less adequate as a characterization of the computational principles employed: cells in V4 and inferotemporal cortex are tuned to complex shapes, object parts, and whole objects, with enormous receptive fields and a qualitatively different form of selectivity [Hubel & Wiesel 1995; Felleman & Van Essen 1991; DiCarlo et al. 2012]."
Section 5.3 (pages 11–12) already establishes this distinction explicitly: V2 does not repeat V1's S-C operation with a wider pooling window but responds to more complex patterns such as angles and curves, while V4 and inferotemporal cortex "operate on completely different computational principles with entirely different selectivity profiles" and "there are no complex cells to model in those regions."
---
**Comments 5:** Figure 1 effectively summarizes the invariance-discrimination tradeoff and provides a strong conceptual anchor for the paper. Similarly, Figures 2 and 3 are clear and helpful in illustrating the distinction between geometry-preserving operations and pooling-based information loss.
**Response 5:** We thank the reviewer for this positive assessment. No changes were made in response to this comment.
---
**Comments 6:** The paper critiques CNNs extensively, but the discussion of modern architectures such as ConvNeXt, hybrid ViT-CNN models, Mamba-based vision architectures, hierarchical transformers, and diffusion-based representations is relatively limited. Including these developments would improve the completeness and timeliness of the review.
**Response 6:** We agree. ConvNeXt and hybrid ViT-CNN models were already discussed in the Vision Transformers subsection of Section 8 (page 19), where their reintroduction of CNN-style inductive biases into ViT architectures is cited as confirmation that local geometric structure was valuable. We have added a new paragraph to Section 8 (pages 19–20) that discusses Mamba-based state-space architectures (Vision Mamba, VMamba) and diffusion-model architectures. The added text reads:
> "State-space model architectures, including Vision Mamba and VMamba, represent a more recent development that replaces self-attention with linear-time sequential scanning, avoiding the quadratic cost of attention while modeling long-range dependencies [Sultan et al. 2026]. Against the three limitations: Mamba-based architectures do not address modality conflation; they largely avoid aggressive spatial subsampling and therefore partially address the pooling problem, though not through the principled combination mechanism the energy model employs; and they do not address the depth problem in the biological sense. Hybrid architectures that combine convolutional feature extraction, state-space scanning, and PDE-based diffusion operations have demonstrated that structural preservation improves when spatial information is explicitly maintained across scales through selective skip fusion rather than discarded by subsampling [Sultan et al. 2026]. Diffusion-model architectures, which typically employ UNet-style encoder-decoder structures with dense skip connections between encoder and decoder stages, represent the field's practical acknowledgment that information discarded during downsampling must be recovered explicitly: the skip connections bypass the pooling-induced information loss rather than preventing it, partially preserving spatial detail for dense prediction tasks."
---
**Comments 7:** The manuscript frequently frames pooling as fundamentally destructive to geometry. While this argument is theoretically motivated, the practical success of deep CNNs in dense prediction tasks such as segmentation, detection, and medical imaging suggests the issue may be more nuanced. The discussion should acknowledge this more explicitly.
**Response 7:** This is addressed in two places in the manuscript. Section 5.3 ("When pooling is defensible, and when it is not," pages 11–12) explicitly distinguishes the first-stage regime, where pooling is a tolerable approximation, from deeper layers where it has no biological warrant, and states: "The problem is not that pooling is a wrong operation but that it is repeated indefinitely past the biological warrant that ends at V3." Section 7 (page 17) has been expanded in this revision with two new sentences that develop the success-condition and failure-boundary analysis specifically for medical imaging and remote sensing:
> "In medical image analysis, CNNs trained and evaluated within a single acquisition protocol achieve competitive performance on classification and segmentation benchmarks, reflecting how well local filter hierarchies capture the texture signatures that distinguish tissue classes within a consistent imaging domain; the same networks generalize poorly when scanners, acquisition protocols, or patient populations change, because protocol-specific texture statistics dominate the learned representations while the anatomical geometry that would support robust generalization was progressively discarded by subsampling. Remote sensing presents an analogous pattern: CNNs perform well on within-distribution land-cover classification but encounter greater difficulty with orientation-invariant object detection and multi-scale geometric reasoning, precisely the tasks where spatial relationships between features, rather than local textural signatures, carry the diagnostic information."
---
**Comments 8:** The section discussing adversarial vulnerability is interesting, but the connection between geometric representation and adversarial robustness remains largely speculative. More direct evidence or references supporting this causal relationship would strengthen the argument.
**Response 8:** The paper presents a mechanistic argument: representations grounded in texture statistics are fragile to targeted perturbations because those statistics can be disrupted while the geometric structure the network largely ignores remains intact. The manuscript is explicit that establishing this causal chain empirically at scale remains an open problem. We have strengthened the passage in Section 6 (page 15) by adding a sentence connecting the prediction to existing empirical findings. The revised sentence reads:
> "This implies that adversarial robustness is not primarily a matter of training procedure or regularization but of the representational basis: networks that represent geometric structure should be more robustly adversarially stable than networks that represent texture statistics, a prediction consistent with the finding that models trained to be shape-biased rather than texture-biased show improved robustness to distribution shift [Geirhos et al. 2018], though establishing the full causal chain empirically at scale remains an open problem."
---
**Comments 9:** The review would benefit from a dedicated section discussing whether current architectural remedies such as skip connections, feature pyramids, attention mechanisms, multi-scale fusion, or positional encoding partially mitigate the identified geometric degradation problem.
**Response 9:** We agree. We have added a paragraph to Section 8 (pages 19–20) covering these as partial architectural mitigations. Skip connections and feature pyramids are discussed directly; attention mechanisms and positional encoding are noted as already covered in the Vision Transformers subsection. The added text reads:
> "Several developments within standard CNN practice have also moved the field in the direction the diagnosis suggests, without explicitly framing themselves as responses to the pooling problem. Many modern architectures reduce or eliminate aggressive max-pooling in favor of strided convolutions, adaptive feature aggregation, or patch-based tokenization, which reduces the Barnard-Casasent loss per stage even if it does not eliminate it. Skip connections in residual architectures allow feature maps from earlier, less-degraded layers to bypass multiple pooling stages and remain available deeper in the network. Feature pyramid networks maintain representations at multiple spatial scales simultaneously, acknowledging that tasks requiring geometric precision depend on information that a purely downsampling hierarchy destroys. These engineering responses are consistent with the paper's diagnosis: the field has found it necessary to reintroduce at various points the spatial information that aggressive pooling discards, without in most cases addressing the underlying mechanism."
---
**Comments 10:** Some arguments against pooling appear to overlook the fact that many modern architectures already reduce or eliminate aggressive pooling operations in favor of strided convolutions, adaptive feature aggregation, or patch-based tokenization. This evolution in the field should be discussed more comprehensively.
**Response 10:** This is addressed in the same new paragraph added to Section 8 (pages 19–20) in response to Comment 9 above, which explicitly notes that strided convolutions, adaptive feature aggregation, and patch-based tokenization have progressively reduced aggressive spatial subsampling in modern architectures, reducing the Barnard-Casasent loss per stage even where they do not eliminate it. This evolution is framed as consistent with, rather than as a refutation of, the paper's diagnosis.
---
**Comments 11:** The manuscript is academically rich, but several paragraphs contain very long sentences that reduce readability. Breaking these into shorter statements would improve clarity and accessibility for a broader readership.
**Response 11:** We respectfully maintain the current sentence structure. The paper is written in a dense academic register in which complex claims are articulated through subordination and relative clauses rather than through sentence fragmentation. This is a deliberate stylistic choice appropriate for a literature review article addressing technically sophisticated readers. Splitting the longer sentences would reduce precision without improving clarity for the intended audience. No changes were made in response to this comment.
---
**Comments 12:** The proposed six desiderata are conceptually meaningful, but they remain relatively abstract. The paper would be strengthened by providing a more concrete architectural framework or preliminary design strategy demonstrating how these principles could realistically be implemented in practice.
**Response 12:** We agree. We have added a closing paragraph to Section 9 (page 23) that maps each desideratum to a concrete architectural commitment and to the closest existing partial implementation. The added text reads:
> "Partial implementations exist for each, indicating that the design space is reachable even if no architecture has yet reached it. D1 is approximated by multi-stream architectures that separate chrominance from luminance before the first learned layer. D2 is most directly approached by PDE-G-CNNs [Smets et al. 2023], which replace subsampling with combination-based geometry extraction at early stages. D3 and D4 are the explicit targets of capsule networks [Sabour et al. 2017; Hinton et al. 2018], which encode pose parameters and propagate them through routing-by-agreement rather than discarding them through pooling, though scaling these mechanisms remains an open engineering challenge. D5 is partially approached by attention mechanisms that concentrate computation on geometrically salient locations rather than distributing it uniformly across the spatial grid. D6 is partially addressed by encoder-decoder architectures with skip connections, which preserve multi-scale spatial information explicitly, and most directly by PDE-G-CNNs, which achieve combination-based integration without irreversible subsampling. The consistent pattern across these partial implementations is that each one works by preserving or reintroducing exactly the spatial information that standard pooling destroys, which is the strongest form of empirical confirmation available for the diagnosis this paper advances."
---
**Comments 13:** DOI: 10.1109/TGRS.2026.3685508 could strengthen the discussion regarding geometry-aware representation learning and structural feature preservation in deep architectures. Additionally, https://doi.org/10.1007/s13244-018-0639-9 may support the discussion on hybrid deep learning frameworks and limitations of conventional CNN feature representations.
**Response 13:** Both references have been evaluated and incorporated. Yamashita et al. (2018, *Insights into Imaging* 9:611–629) has been added to the medical imaging passage in Section 7 (page 17), where it supports the claim about CNN applicability and limitations within consistent acquisition protocols. Sultan et al. (2026, *IEEE Transactions on Geoscience and Remote Sensing* 64:4106218) has been added to the remote sensing passage in Section 7 (page 17) and to the new Mamba and hybrid architectures paragraph in Section 8 (page 19), where DSCH-Net illustrates the performance gains from combining PDE-based diffusion and state-space scanning over standard CNN and Transformer baselines on a structural-preservation task in remote sensing.

Reviewer 2 Report

Comments and Suggestions for Authors

1. The authors use the term "geometric structure" multiple times in the manuscript but fail to provide a precise definition: does it refer to 2D spatial relationships, 3D shape priors, or group isovariability? The evaluation criteria for "biological fidelity" are unclear: does it mean completely mimicking biological mechanisms or merely borrowing core principles? The comparison between the energy model and pooling in Figure 2 is a qualitative illustration, lacking quantitative superposition of response curves. It is recommended that the authors clearly define operational metrics for core terms in the methodology section (e.g., using pose estimation error to measure geometric understanding), supplement with comparative experiments on the response variance of the energy model and max/average pooling under the same input, and discuss the trade-off between "engineering rationality" and "biological realism."

2. The assertion in the manuscript that capsule networks "failed to scale to complex datasets" needs updating (e.g., are there any new developments in 2023-2024?). Section 7, "The Success of CNNs," is relatively short and may weaken the paper's constructive image; some expressions, such as "conceptual conflation," are too strong. It is recommended to add limiting conditions. It is recommended that the authors systematically search for literature related to capsule/G-CNN/PDE-network from 2024-2026 and update the evaluation in Section 8. Section 7 should be expanded to specifically analyze the success conditions and failure boundaries of CNNs in geometrically sensitive tasks such as medical imaging and remote sensing. Some absolute statements should be adjusted to restrictive assertions under the condition of "…" to enhance academic rigor.

3. The manuscript is primarily theoretical, lacking experimental data to support core arguments. For example: The causal chain of "modal confusion leading to texture bias" lacks ablation experiments for verification. Does replacing pooling with an energy model truly improve out-of-distribution generalization? There is no quantitative comparison. The geometric information loss in Figure 3 is illustrative and lacks information-theoretic metrics (such as mutual information decay curves). It is recommended that the authors provide at least a proof-of-concept experiment: comparing the performance of standard CNNs and the "modal separation + energy model" prototype architecture on geometric inference benchmarks such as CLEVR and 3D Shapes. Representative similarity analysis (RSA) is used to quantify the matching degree between different architecture layers and V1/V2/V4 neural responses. Information bottleneck analysis is added to quantify the position information retention rate of each pooling layer.

Author Response

### Response to Reviewer 2
We thank the reviewer for the detailed and substantive comments. We respond to each comment below.
---
**Comments 1:** The authors use the term "geometric structure" multiple times in the manuscript but fail to provide a precise definition: does it refer to 2D spatial relationships, 3D shape priors, or group isovariability? The evaluation criteria for "biological fidelity" are unclear: does it mean completely mimicking biological mechanisms or merely borrowing core principles? The comparison between the energy model and pooling in Figure 2 is a qualitative illustration, lacking quantitative superposition of response curves. It is recommended that the authors clearly define operational metrics for core terms in the methodology section (e.g., using pose estimation error to measure geometric understanding), supplement with comparative experiments on the response variance of the energy model and max/average pooling under the same input, and discuss the trade-off between "engineering rationality" and "biological realism."
**Response 1:** We thank the reviewer for flagging the absence of explicit definitions. Both terms have been defined at their first substantive appearance in the manuscript.
"Biological fidelity" is now defined parenthetically at its first use in Section 1 (page 2), second paragraph:
> "The question of biological fidelity (the degree to which CNN architectural operations implement the same computational logic as their biological counterparts, rather than merely producing similar input-output statistics) in CNNs might seem academic…"
"Geometric structure" is defined at its first substantive use in Section 1 (page 2), third paragraph:
> "The visual cortex exploits the geometric structure of the world as an inductive bias for generalization, where by geometric structure we mean the 2D spatial relationships between oriented contrast features that constitute objects rather than their photometric surface properties such as color or texture…"
A more precise technical elaboration ("the spatial relationships between oriented features, specifically where each edge is relative to each other edge") appears in Section 5.2 (page 11), where the distinction between orientation information and position information is developed in full.
On reflection, we recognize that the abstract, as originally written, may itself have contributed to this reading. The sentence "Pooling is attempting something categorically different" appeared without the qualifying context that pooling is in fact a tolerable first-stage approximation; the problem is its use as a general-purpose mechanism across the full depth of the network. We have revised the abstract accordingly. The updated passage now reads:
> "Pooling is a tolerable first-stage approximation of this behavior, but as a general-purpose invariance mechanism repeated across the full depth of the network it is attempting something categorically different, namely object-level position invariance through spatial subsampling, which achieves its goal by discarding exactly the geometric information that the energy model preserves. Treating pooling as a scalable, indefinitely repeatable implementation of complex cell behavior, rather than as a first-stage approximation with a natural biological endpoint at V3, conflates two operations that differ not in degree but in kind, and crucially it removed the principled criterion for confining the S-C operation to early visual cortex: because pooling was understood as a general-purpose invariance mechanism, the field had no architectural reason to stop repeating it."
We wish to clarify the broader point further. The paper does not argue that pooling should be replaced with the energy model, nor does it frame the central issue as a trade-off between engineering rationality and biological realism. The argument is more specific: pooling is a rough but tolerable approximation of complex cell behavior at the first stage of the network, where the simple-to-complex cell transformation has genuine biological warrant; the problem is that this operation is repeated indefinitely past V3, into regions of the hierarchy that have no complex cells to model and where the biological system performs qualitatively different computations. Both points are developed in Section 5.3 (pages 11–12), which states explicitly: "The problem is not that pooling is a wrong operation but that it is repeated indefinitely past the biological warrant that ends at V3 and that what it is being used to approximate, complex cell behavior, was never a position invariance mechanism to begin with." The energy model serves in this paper not as a proposed replacement for pooling but as a diagnostic instrument. This is stated explicitly in Section 5.4 (page 12): "The paper does not claim that every pooling stage should have been replaced with an energy model computation, or that implementing the energy model throughout the network would resolve the limitations diagnosed here." We respectfully suggest that the trade-off framing the reviewer proposes is not the one the paper is making.
Regarding the request to supplement Figure 2 (page 10) with quantitative response curves: Figure 2 is a schematic illustration of the computational distinction between the two operations, which is the appropriate form for a figure in a literature review article making an architectural argument rather than reporting a new measurement. The relevant quantitative analysis of energy model and pooling responses is available in the primary literature we cite, and reproducing it here would not strengthen the architectural argument the figure is making.
---
**Comments 2:** The assertion in the manuscript that capsule networks "failed to scale to complex datasets" needs updating (e.g., are there any new developments in 2023-2024?). Section 7, "The Success of CNNs," is relatively short and may weaken the paper's constructive image; some expressions, such as "conceptual conflation," are too strong. It is recommended to add limiting conditions. It is recommended that the authors systematically search for literature related to capsule/G-CNN/PDE-network from 2024-2026 and update the evaluation in Section 8. Section 7 should be expanded to specifically analyze the success conditions and failure boundaries of CNNs in geometrically sensitive tasks such as medical imaging and remote sensing. Some absolute statements should be adjusted to restrictive assertions under the condition of "…" to enhance academic rigor.
**Response 2:** We have made the following changes in response to this comment.
The capsule network scaling claim in Section 8.1 (page 18) has been softened. The revised sentence reads:
> "The practical limitation is significant: routing-by-agreement is computationally expensive and capsule networks have shown considerably more difficulty scaling to complex natural image datasets than standard CNNs, with subsequent architectural variants improving parameter efficiency and domain-specific performance without resolving the core scalability constraints, suggesting that the approach of learning the geometric relationship that the energy model achieves by construction faces efficiency challenges that remain open."
Section 7 (page 17) has been expanded with two sentences that develop the success-condition and failure-boundary analysis specifically for medical imaging and remote sensing:
> "In medical image analysis, CNNs trained and evaluated within a single acquisition protocol achieve competitive performance on classification and segmentation benchmarks, reflecting how well local filter hierarchies capture the texture signatures that distinguish tissue classes within a consistent imaging domain; the same networks generalize poorly when scanners, acquisition protocols, or patient populations change, because protocol-specific texture statistics dominate the learned representations while the anatomical geometry that would support robust generalization was progressively discarded by subsampling. Remote sensing presents an analogous pattern: CNNs perform well on within-distribution land-cover classification but encounter greater difficulty with orientation-invariant object detection and multi-scale geometric reasoning, precisely the tasks where spatial relationships between features, rather than local textural signatures, carry the diagnostic information."
We have added a sentence to each of the G-CNN subsection (Section 8.2, page 18) and PDE-G-CNN subsection (Section 8.3, page 18–19) acknowledging recent architectural extensions. In Section 8.2:
> "Recent extensions including attention-based G-convolutions and separable G-convolutions have improved parameter efficiency and retained equivariance at greater scale, but the core constraint remains: enriching early-layer geometry without changing the subsampling mechanism leaves the compounding geometric loss intact."
In Section 8.3:
> "Recent work has extended the framework with axiomatic derivations and additional architectural variants, but competitive performance with standard CNNs and ViTs on large benchmarks remains an open challenge."
Regarding "conceptual conflation": we respectfully maintain this term. It is not used rhetorically but as a precise description of a specific category error, namely treating two operations that differ in kind (geometry extraction by combination versus position invariance by subsampling) as approximations of one another. The paper defines the conflation explicitly and develops its consequences across two sections; the term is not an absolute judgment but a diagnosis of a specific architectural decision, and we do not believe a qualifying condition would make it more precise.
---
**Comments 3:** The manuscript is primarily theoretical, lacking experimental data to support core arguments. For example: The causal chain of "modal confusion leading to texture bias" lacks ablation experiments for verification. Does replacing pooling with an energy model truly improve out-of-distribution generalization? There is no quantitative comparison. The geometric information loss in Figure 3 is illustrative and lacks information-theoretic metrics (such as mutual information decay curves). It is recommended that the authors provide at least a proof-of-concept experiment: comparing the performance of standard CNNs and the "modal separation + energy model" prototype architecture on geometric inference benchmarks such as CLEVR and 3D Shapes. Representative similarity analysis (RSA) is used to quantify the matching degree between different architecture layers and V1/V2/V4 neural responses. Information bottleneck analysis is added to quantify the position information retention rate of each pooling layer.
**Response 3:** The claims the reviewer identifies as lacking experimental support are in fact supported by extensive cited empirical literature. The connection between modality conflation and texture bias is documented by Geirhos et al. (2018, 2019), Baker et al. (2018), and Hermann et al. (2019), each of which provides direct empirical evidence for the causal relationship the reviewer asks us to verify with ablations. The out-of-distribution fragility of CNNs is documented by Arjovsky (2021) and Krueger et al. (2021) among others. These are not novel theoretical claims requiring new verification; they are documented phenomena whose structural origin this literature review diagnoses. The contribution of this paper is to synthesize and analyze these existing findings within a unified architectural framework, not to replicate the underlying measurements.
The reviewer asks whether replacing pooling with an energy model would improve out-of-distribution generalization. As noted in our response to Comment 1, the paper does not make this claim. Section 5.4 (page 12) states explicitly: "The paper does not claim that every pooling stage should have been replaced with an energy model computation, or that implementing the energy model throughout the network would resolve the limitations diagnosed here." There is therefore no quantitative comparison to provide, because the comparison the reviewer requests is one the paper does not advocate.
Regarding Figure 3 (page 14): it is a schematic illustration of the conceptual argument about positional information loss, not a measurement. The addition of mutual information decay curves or information bottleneck analysis across pooling layers would constitute an independent empirical contribution requiring its own experimental design, validation, and analysis. Similarly, proof-of-concept experiments on CLEVR or 3D Shapes and representational similarity analysis against V1/V2/V4 responses are each substantial research programs in their own right. The reviewer's requested additions, taken together, would constitute several independent papers' worth of original empirical work rather than additions to a literature review article. We note that where the paper makes forward-looking empirical claims, for instance regarding adversarial robustness, it explicitly flags them as open problems. Section 6 (page 15) states: "establishing this empirically at scale remains an open problem," which we consider an appropriate and honest characterization of the state of the field.

Reviewer 3 Report

Comments and Suggestions for Authors
  1. The paper contains no new experiments, datasets, or original formal results; the figures are schematic illustrations rather than measured outcomes (Figure 3 in particular reads as a hand-constructed demonstration, not data). You should label this as a perspective/review.
  2. Do not group max pooling, average pooling, and strided convolution as equivalent. You repeatedly state these "discard spatial information in the same fundamental way" (lines 384, 509–510). Average pooling and strided convolution are linear operations whose loss is primarily aliasing, not selection-and-discard. There is an active and directly relevant literature on this — anti-aliased downsampling / BlurPool, and analyses of why CNNs are not shift-invariant — that you neither cite nor engage. Its absence is conspicuous given your topic, and engaging it would both strengthen and appropriately qualify your argument.
  3. The claims that the Hubel–Wiesel framework "stops applying altogether" beyond V3 (lines 167–168) and that the hierarchy is a clean feedforward succession of qualitatively different stages omit recurrence, feedback, and the genuinely fuzzy functional boundaries between areas. The strict V1→V2→V3→V4→IT staging is a reasonable idealization but is presented as settled on the basis of two textbook-level references [20,21]. Add hedging and stronger primary sources.
  4. Figure 5 (Section 8) uses the D1–D6 labels and short names that aren't defined until Section 9. Either move the desiderata earlier or add a forward pointer in the caption so the table is interpretable where it appears.
  5. Recheck multi-reference clusters against the specific claim. Several cite sources that don't support the adjacent statement. Examples: "the spatiotemporal energy model of Adelson and Bergen [2,18,19]" (line ~357) — only [19] is Adelson & Bergen; [2] is Hubel–Wiesel 1968 and [18] is Schiller et al. The V1 orientation-column claim cites "[28,52–54]" (line ~670), but [28] (Barnard–Casasent, shift invariance) has nothing to do with orientation-column physiology. The PDE-G-CNN description cites "[19,28,55]" (line ~684) where [19] and [28] look misplaced relative to [55]. Audit every cluster.
  6. "Second-order overfitting" (line 530) is non-standard and undefined. Define it precisely or use established terminology.
  7. Figure 4 appears to use a photograph of a real, identifiable person (an astronaut portrait) relabeled "sea lion." Confirm licensing and that you generated the example yourself; if it's adapted from another source, cite it. A documented ImageNet sample would avoid the issue.

Author Response

### Response to Reviewer 3
We thank the reviewer for the thorough and precise reading of the manuscript, particularly the close attention to citation accuracy and figure licensing. We respond to each comment below.
---
**Comments 1:** The paper contains no new experiments, datasets, or original formal results; figures are schematic illustrations rather than measured outcomes (Figure 3 in particular reads as hand-constructed). The reviewer suggests labeling it as a perspective/review. Note: The article type is set on the submission platform, not in the LaTeX, as the MDPI class in use does not support `\articletype{}`, and the editor already knows the submission category.
**Response 1:** We thank the reviewer for this observation. We clarify that the manuscript is best characterized as a review article with a perspective orientation toward future directions. It surveys the biological foundations of CNNs and the existing empirical and theoretical literature on their limitations before advancing a forward-looking architectural argument. The article type was set accordingly on the submission platform. No manuscript change was required, as the review character of the work is already reflected in the structure, scope, and citation practice of the paper.
---
**Comments 2:** The manuscript repeatedly states that max pooling, average pooling, and strided convolution "discard spatial information in the same fundamental way." This is incorrect: average pooling and strided convolution are linear operations whose primary loss is aliasing, not selection-and-discard. The relevant literature on anti-aliased downsampling (BlurPool) and CNN shift-invariance is neither cited nor engaged, which is conspicuous given the paper's topic.
**Response 2:** We agree with this correction. The sentence in Section 6 (page 13), first paragraph, has been revised to acknowledge the mechanistic distinction between pooling operations. The core argument, that all three operations discard positional information irreversibly, is preserved, but the mechanisms are now distinguished. We have added a citation to Zhang (2019), "Making Convolutional Networks Shift-Invariant Again" (ICML 2019), which provides the directly relevant analysis of aliasing in downsampling operations. The revised sentence reads:
> "Max pooling, average pooling, and strided convolution achieve position tolerance through spatial subsampling by different mechanisms, with max pooling operating by selection-and-discard and average pooling and strided convolution by linear aggregation whose primary failure mode is aliasing rather than selection [Zhang 2019], yet all three discard positional information irreversibly, which is the property that matters for the argument advanced here."
---
**Comments 3:** The claim that the Hubel–Wiesel framework "stops applying altogether" beyond V3 and the presentation of a clean feedforward V1→V2→V3→V4→IT staging omit recurrence, feedback, and the genuinely fuzzy functional boundaries between cortical areas. The two textbook-level references are insufficient. Stronger primary sources and explicit hedging are required.
**Response 3:** We partially agree with this comment. The phrase "stops applying altogether" was overstated and has been revised. The updated text appears in Section 2 (page 4), in the paragraph discussing the S-C organizational principle. The revised passage reads:
> "…beyond V3 the Hubel-Wiesel framework becomes progressively less adequate as a characterization of the computational principles employed: cells in V4 and inferotemporal cortex are tuned to complex shapes, object parts, and whole objects, with enormous receptive fields and a qualitatively different form of selectivity [Hubel & Wiesel 1995; Felleman & Van Essen 1991; DiCarlo, Zoccolan & Rust 2012]."
DiCarlo, Zoccolan & Rust (2012, *Neuron* 73:415–434) has been added as a primary source for the V4/IT selectivity claim.
We respectfully disagree, however, with the characterization that the paper presents the hierarchy as a clean feedforward succession that omits recurrence and feedback. The paper's argument runs in precisely the opposite direction: Section 2 states explicitly that "each stage in the real visual hierarchy does something qualitatively different from the stage below it," and this biological complexity is the standard against which CNN design is found wanting. The paper uses the non-repetitive character of the biological hierarchy as its central argument against the CNN's uniform S-C stack, and does not endorse a simplified view of cortical processing.
---
**Comments 4:** Figure 5 (Section 8) uses the D1–D6 labels and short names that are not defined until Section 9. Either move the desiderata presentation earlier or add a forward pointer in the Figure 5 caption so the table is interpretable at the point where it appears.
**Response 4:** We agree. The Figure 5 caption (page 20) has been updated to include a forward pointer. The revised caption now reads:
> "Coverage of six proposed desiderata (D1–D6, defined in Section 9) by major CNN alternatives…"
---
**Comments 5:** Several multi-citation clusters cite sources that do not support the adjacent claim. Specific examples flagged: (i) "the spatiotemporal energy model of Adelson and Bergen [2,18,19]" — only [19] is Adelson & Bergen; (ii) V1 orientation-column claim cites "[28,52–54]" — [28] (Barnard–Casasent, shift invariance) has nothing to do with orientation-column physiology; (iii) PDE-G-CNN description cites "[19,28,55]" — [19] and [28] appear misplaced. Action required: audit every multi-reference cluster in the manuscript.
**Response 5:** We thank the reviewer for the careful reading. All multi-reference clusters in the manuscript were audited and the following corrections were made throughout the manuscript:
1. The sentence describing "the spatiotemporal energy model of Adelson and Bergen" has been restructured to attribute the energy model to Adelson & Bergen alone, with Hubel & Wiesel (1968) and Schiller et al. (1976) cited separately as the physiological characterizations of complex cell behavior that the model formalizes.
2. The V1 orientation-column claim previously included Barnard & Casasent (1990) among its citations; this reference concerns shift invariance in the Neocognitron and has no bearing on orientation column physiology. It has been removed and replaced with Hubel & Wiesel (1962, 1968) as the appropriate primary sources.
3. The description of PDE-G-CNNs replacing the convolution-pooling-ReLU trifecta with PDE solvers on Lie groups previously cited Barnard & Casasent and Adelson & Bergen alongside Smets et al. These references appear correctly in the subsequent sentence where the explicit comparison to the energy model and shift invariance is made; they have been removed from the earlier sentence, leaving only Smets et al. (2023).
4. Desiderata 2 and 6 previously cited Felleman & Van Essen (1991) in clusters supporting claims about geometry extraction and information-preserving combination. As Felleman & Van Essen concerns hierarchical connectivity mapping rather than the energy model or positional information loss, it has been removed from both clusters.
---
**Comments 6:** "Second-order overfitting" (line 530) is non-standard terminology. Must be precisely defined or replaced with established terminology.
**Response 6:** We agree that this term required definition at point of use. The passage in Section 6 (page 14) has been revised to include an inline definition. The revised passage reads:
> "The CNN has become closely fitted to the statistics of the training distribution, a form of second-order overfitting, specifically overfitting to the statistical regularities of the training distribution rather than to individual training labels, a failure mode that maps directly onto out-of-distribution generalization problems [Sa-Couto & Wichert 2021], that partially resists correction because the geometric structure that would support distribution-robust generalization was progressively discarded during training…"
The term was introduced in Sa-Couto & Wichert (2021), "Simple Convolutional-Based Models: Are They Learning the Task or the Data?", which is already cited in the preceding sentence; the revised passage makes the origin and meaning of the term explicit.
---
**Comments 7:** Figure 4 appears to use a photograph of a real, identifiable person (an astronaut portrait) relabeled "sea lion." Licensing must be confirmed and the example must have been generated by the authors. If adapted from another source, it must be cited. Using a documented ImageNet sample would avoid the issue entirely.
**Response 7:** We thank the reviewer for raising this concern. The original adversarial example figure used a scikit-image built-in photograph of an identifiable person. This has been replaced with a CC0 1.0 Public Domain photograph of a giant panda (Aibao, photographed by Cheng Shiyi, Wikimedia Commons). The figure caption (Figure 4, page 16) now reads:
> "Illustrative adversarial example following Goodfellow et al. [2014]. The original image (A) is correctly classified with high confidence. Adding an imperceptible noise pattern (B), amplified 10× for visibility, produces an image indistinguishable from the original to human observers (C), yet the network misclassifies it with high confidence. This fragility is not random: it is the predictable consequence of representations grounded in local texture statistics rather than geometric structure. A representation that internalized the geometry of visual categories would be robust by construction, because geometric structure is stable under small perturbations. Panda photograph by Cheng Shiyi, CC0 1.0 Public Domain (Wikimedia Commons); confidence scores are illustrative."
This resolves both the licensing and attribution concerns raised.

Round 2

Reviewer 1 Report

Comments and Suggestions for Authors

All my concerns have been addressed significantly. No more comments. 

Reviewer 3 Report

Comments and Suggestions for Authors

no comments

Back to TopTop