GPU-Parallelization of the Numerical Solution of Large-Scale Transient Partial Differential Equations
Abstract
1. Introduction
2. Mathematical Background, the Numerical Methods Being Examined
2.1. The Heat Transfer or Diffusion Equation, Spatial, and Time-Domain Discretization
- U is the temperature [°C, K] or—in the case of diffusion—concentration (mass [kg/m2] or molar [mol/m2] concentration);
- t is time [s];
- is the position vector [m];
- D is the diffusion coefficient [m2/s]; and
- is the Laplacian operator.
2.2. Handling Boundary Conditions
- Calculating the actual boundary values using the formula of the analytical solution at each timestep;
- Pre-calculating all the used boundary values before the loop for time marching starts.
2.3. The Algorithms Under Investigation
- Euler’s explicit method;
- Constant Neighbours (CNe) method;
- CpC method.
2.3.1. Euler’s Explicit Method
2.3.2. The Constant Neighbor Method
2.3.3. The CpC Method
2.4. GPU Parallelization and Implementation Strategy
2.4.1. Parallelization of Euler and CNe
| Listing 1. Kernel of Euler, CNe, and the 1st stage of CpC. This is a verbatim quotation from [21]. |
| #pragma OPENCLEXTENSION cl_khr_fp64: enable __kernel void calc_cell_2D( __global const double *before, __global double *after, double2 coeffs, uint2 size, long4 neighbours) { const size_t index = get_global_id(0) + size.x * get_global_id(1); __global const double *old_cell = before + index; after[index] = (1 − 2 * (coeffs.x + coeffs.y)) * *old_cell + coeffs.x * (old_cell[neighbours.even.x] + old_cell[neighbours.odd.x]) + coeffs.y * (old_cell[neighbours.even.y] + old_cell[neighbours.odd.y]); } |
- before, after are pointers to the array containing data before and after the timestep;
- coeffs is the [coeffx, coeffy] vector containing numerical coefficients for computation;
- size is a vector containing the [Nx, Ny] number of data points; and
- neighbours is a vector containing pointer differences between the current datapoint and its neighbors.
| Listing 2. Kernel to apply boundary data. |
| #pragma OPENCLEXTENSION cl_khr_fp64: enable __kernel void load_bound( __global double *datagrid, __global const double *bound_evol, const uint tm_offs, const unsigned long *idx_tbl, uint bndsize) { __global const double *timeframe = bound_evol + tm_offs; const size_t i = get_global_id(0); datagrid[idx_tbl[i]] = timeframe[i]; } |
- datagrid are pointers to the datagrid the boundary data is being loaded into;
- bound_evol is a pointer to the array containing the boundary evolution data;
- tm_offs is the time-proportional offset inside array bound_evol;
- idx_tbl is an index-table containing the offset values of the boundary data-points inside datagrid; and
- bndsize is the size of the boundary values per timestep (same as the total number of work-items).
- Offset values of the 4 corner nodes;
- The Nx − 2 offset values of the j = 1 vertex;
- The Ny − 2 offset values of the i = Nx vertex;
- The Nx − 2 offset values of the j = Ny vertex;
- The Ny − 2 offset values of the i = 1 vertex.
- Load initial and boundary data from file to host RAM;
- Initialize OpenCL and build kernel code;
- Allocate GPU memory for 2 full data-grids, containing a copy of the initial data, and one more GPU RAM array for boundary data (with “using HOST pointer” mode);
- Acquire kernel instances and parametrize them for the time evolution loop;
- Enqueue execution of boundary loader kernel;
- Enqueue execution of inner-domain kernel function. Roles of the 2 full-grid memory buffers depend on the parity of the current timestep;
- Swap the roles of the 2 full-grid memory buffers to prepare for next timestep iteration;
- Go back to 5 and repeat the loop until the last timestep is handled;
- Enqueue reading back the result from the last written full-grid GPU buffer to RAM;
- Wait until all data movements are finished;
- Release kernel resources and GPU RAM;
- Write results to file;
- On exit, release any remaining OpenCL resources.
2.4.2. Parallelization of CpC
| Listing 3. Kernel of the 2nd stage of CpC. This is a verbatim quotation from [21]. |
| __kernel void stage2p05_calc_cell_2D( __global const double *stage0, __global const double *stage1, __global double *after, double2 coeffs, uint2 size, long4 neighbours) { const size_t index = get_global_id(0) + size.x * get_global_id(1); __global const double *old_cell = stage0 + index; __global const double *midpoint = stage1 + index; after[index] = (1 − 2 * (coeffs.x + coeffs.y)) * *old_cell + coeffs.x * (midpoint[neighbours.even.x] + midpoint[neighbours.odd.x]) + coeffs.y * (midpoint[neighbours.even.y] + midpoint[neighbours.odd.y]); } |
- stage0 is a pointer to the array containing data before stage 1;
- stage1 is a pointer to the array containing the result of stage 1;
- after is a pointer to the array containing data after all the stages (often the same as stage0);
- coeffs is the [coeffx, coeffy] vector containing numerical coefficients for computation;
- size is a vector containing the [Nx, Ny] number of data points; and
- neighbours is a vector containing pointer differences between the current datapoint and its neighbours.
- Load initial and boundary data from file to CPU RAM;
- Initialize OpenCL and build kernel code;
- Allocate GPU memory for 2 full datagrids, containing a copy of the initial data, and one more GPU RAM array for boundary data (with “using HOST pointer” mode).
- Acquire kernel instances and parametrize them for the time evolution loop;
- Enqueue execution of boundary loader kernel for the mid-point buffer;
- Enqueue execution of stage 1 inner-domain kernel function;
- Enqueue execution of boundary loader kernel for the full-step buffer;
- Enqueue execution of stage 2 inner-domain kernel function;
- Go back to 5 and repeat the loop until the last timestep is handled;
- Enqueue reading back result from full-step GPU buffer to RAM;
- Wait until all data movements are finished;
- Release kernel resources and GPU RAM;
- Write results to file;
- On exit, release any remaining OpenCL resources.
3. General Information About the Way of Investigation
3.1. The Types of Examinations Being Performed
- General benchmarking of the three methods: The algorithmic error is shown as a function of the computation time, as well as that of the timestep size. Curves belonging to different algorithms are plotted on the same figure. The execution time is examined as a function of the total number of spatial grid points. This allows seeing the trade-off between computations with different numbers of spatial dimensions;
- Examining the depletion of GPU, as the system-size is being enlarged: The computation time is plotted both as a function of the total number of spatial grid-points and—for the CpC method—as a function of the timestep count;
- Examining the error convergence of the CpC method when refining the timescale: The algorithmic error is examined as a function of the timestep size. The error should converge to a small residual value, as the timestep size is decreased. Curves for multiple grids are presented on the same figure.
3.2. Initial Value Problems
3.2.1. One-Dimensional Initial Value Problems
- is the norming factor, and
- D = 1 is the diffusion coefficient.
3.2.2. Two-Dimensional Initial Value Problem
- is the new norming factor, and
- D = 1 is the diffusion coefficient.
3.3. Measuring Execution Time, Definition of Algorithmic Error
3.4. The Technical Specifications of the Computer Being Used
- CPU: Gen11 Intel Core i7-11700F, 2.5 GHz;
- RAM: 64 GB;
- OS: Windows 11 Enterprise (24H2), ×64;
- GPU: NVIDIA, 12GB;
- OpenCL: Intel v2020.3.494;
- MATLAB: R2025B.
4. General Benchmarking of the Three Methods
- The first type of plot shows algorithmic error as a function of the execution time. For benchmarking reasons, data of different algorithms are plotted as different curves on the same figure. For different numbers of spatial dimensions, different figures are created;
- The second type of plot shows the algorithmic error as a function of the timestep size, to enable parameter-wise benchmarking.
4.1. Scaling Parameters
4.2. Results
4.2.1. Error—Computation Time Characteristics, One-Dimensional Computation
4.2.2. Error—Computation Time Characteristics, Two-Dimensional Computation
4.3. Discussion
5. Acceleration and Depletion of Parallel Computation Capacity for Increasing System Size
- The data acquisitions of Section 4 are repeated with a grid-sweep up to a much higher—40,000 by 40,000—node count. Computation time is plotted as a function of the total number of spatial nodes;
- In the case of the CpC method, the above experiment is repeated with two more timescales, and another set of timescales and grids (with a much lower node count) are introduced to gain insight into the dependency on both the node count and the timestep count. As no accurate fitting model was found by the authors to explain the results, two types of figures were plotted to analyze performance on huge systems:
- ○
- one to show the effect of timestep count, and
- ○
- one to show the dependence on the spatial resolution.
5.1. Scaling Parameters
5.1.1. Comparing All the Methods on Big Systems
5.1.2. Analyzing Performance of CpC Method
5.2. Results
5.2.1. Computation Time of CpC, Depending on Timestep Count
5.2.2. Computation Time of CpC, Depending on Spatial Node Count, Seeing Depletion of GPU
5.2.3. Comparing Different Algorithms, Seeing Depletion of GPU
5.3. Discussion
6. Error Convergence of CpC Method, as Timescale Is Being Refined
6.1. Scaling Parameters
6.2. Results
6.3. Discussion
7. Summary and Conclusions
Author Contributions
Funding
Data Availability Statement
Acknowledgments
Conflicts of Interest
References
- Ledford, H. “RAMmageddon” hits labs: AI-driven memory shortage is impacting science. Nature 2026, 651, 861–862. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Hundsdorfer, W.; Verwer, J. Numerical Solution of Time-Dependent Advection-Diffusion-Reaction Equations; Springer Series in Computational Mathematics; Springer: Berlin/Heidelberg, Germany, 2003; Volume 33. [Google Scholar] [CrossRef] [Scilit]
- Beuken, L.; Cheffert, O.; Tutueva, A.; Butusov, D.; Legat, V. Numerical Stability and Performance of Semi-Explicit and Semi-Implicit Predictor–Corrector Methods. Mathematics 2022, 10, 2015. [Google Scholar] [CrossRef] [Scilit]
- Essongue, S.; Ledoux, Y.; Ballu, A. Speeding up mesoscale thermal simulations of powder bed additive manufacturing thanks to the forward Euler time-integration scheme: A critical assessment. Finite Elem. Anal. Des. 2022, 211, 103825. [Google Scholar] [CrossRef] [Scilit]
- Gagliardi, F.; Moreto, M.; Olivieri, M.; Valero, M. The international race towards Exascale in Europe. CCF Trans. High Perform. Comput. 2019, 1, 3–13. [Google Scholar] [CrossRef] [Scilit]
- Reguly, I.Z.; Mudalige, G.R. Productivity, performance, and portability for computational fluid dynamics applications. Comput. Fluids 2020, 199, 104425. [Google Scholar] [CrossRef] [Scilit]
- Ali, N.H.M.; Abdullah, R.; Lee, K.J. A comparative study of explicit group iterative solvers on a cluster of workstations. Parallel Algorithms Appl. 2004, 19, 237–255. [Google Scholar] [CrossRef] [Scilit]
- Ewedafe, S.U.; Shariffudin, R.H. On the Parallel Design and Analysis for 3-D ADI Telegraph Problem with MPI. Int. J. Adv. Comput. Sci. Appl. 2014, 5, 18. [Google Scholar] [CrossRef] [Scilit]
- Temirbekov, A.; Altybay, A.; Temirbekova, L.; Kasenov, S. Development of parallel implementation for the Navier-Stokes equation in doubly connected areas using the fictitious domain method. East.-Eur. J. Enterp. Technol. 2022, 2, 38–46. [Google Scholar] [CrossRef] [Scilit]
- Xie, J.; He, J.; Bao, Y.; Chen, X. A low-communication-overhead parallel method for the 3D incompressible Navier-Stokes equations. arXiv 2021, arXiv:2104.08863. [Google Scholar] [CrossRef] [Scilit]
- Colmenares, J.; Galizia, A.; Ortiz, J.; Clematis, A.; Rocchia, W. A Combined MPI-CUDA Parallel Solution of Linear and Nonlinear Poisson-Boltzmann Equation. BioMed Res. Int. 2014, 2014, 560987. [Google Scholar] [CrossRef] [Scilit] [PubMed]
- Rueda Castillo, D.; Kaya, U.; Richter, T. An explicit time integration method for Boussinesq approximation. Proc. Appl. Math. Mech. 2024, 24, e202400050. [Google Scholar] [CrossRef] [Scilit]
- Tan, L.; Huang, M.; Ying, W. A GPU-accelerated Cartesian grid method is proposed for solving the heat, wave, and Schrodinger equations on irregular domains. arXiv 2024. [Google Scholar] [CrossRef] [Scilit]
- De Luca, P.; Fiorillo, G.; Marcellino, L. A GPU-CUDA Numerical Algorithm for Solving a Biological Model. AppliedMath 2025, 5, 178. [Google Scholar] [CrossRef] [Scilit]
- Łach, Ł.; Svyetlichnyy, D. Advances in Numerical Modeling for Heat Transfer and Thermal Management: A Review of Computational Approaches and Environmental Impacts. Energies 2025, 18, 1302. [Google Scholar] [CrossRef] [Scilit]
- Bisbas, G.; Nelson, R.; Louboutin, M.; Luporini, F.; Kelly, P.H.J.; Gorman, G. Automated MPI-X Code Generation for Scalable Finite-Difference Solvers. In Proceedings of the 2025 IEEE International Parallel and Distributed Processing Symposium (IPDPS), Milano, Italy, 3–7 June 2025; pp. 689–701. [Google Scholar] [CrossRef] [Scilit]
- Essongue, S.; Diarra, B.; Lacoste, E. Runge–Kutta–Chebyshev schemes to accelerate thermal modelling of additive manufacturing processes. Finite Elem. Anal. Des. 2026, 256, 104527. [Google Scholar] [CrossRef] [Scilit]
- Kareem Jalghaf, H.; Omle, I.; Kovács, E. A Comparative Study of Explicit and Stable Time Integration Schemes for Heat Conduction in an Insulated Wall. Buildings 2022, 12, 824. [Google Scholar] [CrossRef] [Scilit]
- Kovács, E. A class of new stable, explicit methods to solve the non-stationary heat equation. Numer. Methods Partial Differ. Equ. 2021, 37, 2469–2489. [Google Scholar] [CrossRef] [Scilit]
- Kovács, E.; Nagy, Á.; Saleh, M. A set of new stable, explicit, second order schemes for the non-stationary heat conduction equation. Mathematics 2021, 9, 2284. [Google Scholar] [CrossRef] [Scilit]
- Koics, D.; Kovács, E.; Hornyák, O. Effects of OpenCL-Based Parallelization Methods on Explicit Numerical Methods to Solve the Heat Equation. Computers 2024, 13, 250. [Google Scholar] [CrossRef] [Scilit]
- Mátyás, L.; Barna, I.F. General Self-Similar Solutions of Diffusion Equation and Related Constructions. Rom. J. Phys. 2022, 67, 101. [Google Scholar]
- Garvey, J.D.; Abdelrahman, T.S. A Strategy for Automatic Performance Tuning of Stencil Computations on GPUs. Sci. Program. 2018, 2018, 6093054. [Google Scholar] [CrossRef] [Scilit]
- OpenCL. Wikipedia. 18 July 2026. Available online: https://en.wikipedia.org/w/index.php?title=OpenCL&oldid=1364852294 (accessed on 7 September 2026).
- OpenCL. The Open Standard for Parallel Programming of Heterogeneous Systems. The Khronos Group. Available online: https://www.khronos.org/opencl/ (accessed on 7 September 2026).
- Halbiniak, K.; Szustak, L.; Olas, T.; Wyrzykowski, R.; Gepner, P. Exploration of OpenCL Heterogeneous Programming for Porting Solidification Modeling to CPU-GPU Platforms. Concurr. Comput. Pract. Exp. 2021, 33, e6011. [Google Scholar] [CrossRef] [Scilit]


























| Nt | Δt | Nt | Δt |
|---|---|---|---|
| 9 | 1.111 × 10−1 | 15,000 | 6.667 × 10−5 |
| 15 | 6.667 × 10−2 | 30,000 | 3.333 × 10−5 |
| 30 | 3.333 × 10−2 | 50,000 | 2.000 × 10−5 |
| 50 | 2.000 × 10−2 | 90,000 | 1.111 × 10−5 |
| 150 | 6.667 × 10−3 | 150,000 | 6.667 × 10−6 |
| 300 | 3.333 × 10−3 | 300,000 | 3.333 × 10−6 |
| 500 | 2.000 × 10−3 | 500,000 | 2.000 × 10−6 |
| 900 | 1.111 × 10−3 | 900,000 | 1.111 × 10−6 |
| 1500 | 6.667 × 10−4 | 1,500,000 | 6.667 × 10−7 |
| 3000 | 3.333 × 10−4 | 3,000,000 | 3.333 × 10−7 |
| 5000 | 2.000 × 10−4 | 5,000,000 | 2.000 × 10−7 |
| 9000 | 1.111 × 10−4 | 9,000,000 | 1.111 × 10−7 |
| Nx | Δx |
|---|---|
| 100 | 1.010 × 10−1 |
| 1000 | 1.001 × 10−2 |
| Nt | Δt | Nt | Δt |
|---|---|---|---|
| 9 | 1.111 × 10−2 | 5000 | 2.000 × 10−5 |
| 15 | 6.667 × 10−3 | 9000 | 1.111 × 10−5 |
| 30 | 3.333 × 10−3 | 15,000 | 6.667 × 10−6 |
| 50 | 2.000 × 10−3 | 30,000 | 3.333 × 10−6 |
| 150 | 6.667 × 10−4 | 50,000 | 2.000 × 10−6 |
| 300 | 3.333 × 10−4 | 90,000 | 1.111 × 10−6 |
| 500 | 2.000 × 10−4 | 150,000 | 6.667 × 10−7 |
| 900 | 1.111 × 10−4 | 300,000 | 3.333 × 10−7 |
| 1500 | 6.667 × 10−5 | 600,000 | 1.667 × 10−7 |
| 3000 | 3.333 × 10−5 | 900,000 | 1.111 × 10−7 |
| Nx | Ny | Nx * Ny | Δx⋅Δy |
|---|---|---|---|
| 40 | 40 | 1600 | 6.575 × 10−4 |
| 100 | 100 | 10,000 | 1.020 × 10−4 |
| Tinit | Tfin | Nt | Δt |
|---|---|---|---|
| 1 | 2 | 500 | 2.000 × 10−3 |
| Nx | Δx | Nx | Δx |
|---|---|---|---|
| 40 | 2.564 × 10−1 | 4,000,000 | 2500 × 10−6 |
| 120 | 8.403 × 10−2 | * 5,000,000 | 2000 × 10−6 |
| 400 | 2.506 × 10−2 | * 7,000,000 | 1429 × 10−6 |
| 1200 | 8.340 × 10−3 | * 9,000,000 | 1111 × 10−6 |
| 4000 | 2.501 × 10−3 | 12,000,000 | 8333 × 10−7 |
| 12,000 | 8334 × 10−4 | * 16,000,000 | 6250 × 10−7 |
| 40,000 | 2.500 × 10−4 | * 22,000,000 | 4545 × 10−7 |
| * 50,000 | 2.000 × 10−4 | * 30,000,000 | 3333 × 10−7 |
| * 70,000 | 1.429 × 10−4 | 40,000,000 | 2500 × 10−7 |
| * 90,000 | 1.111 × 10−4 | * 50,000,000 | 2000 × 10−7 |
| 120,000 | 8.333 × 10−5 | * 70,000,000 | 1429 × 10−7 |
| * 160,000 | 6.250 × 10−5 | * 90,000,000 | 1111 × 10−7 |
| * 220,000 | 4.545 × 10−5 | 120,000,000 | 8333 × 10−8 |
| * 300,000 | 3.333 × 10−5 | * 160,000,000 | 6250 × 10−8 |
| 400,000 | 2.500 × 10−5 | * 220,000,000 | 4545 × 10−8 |
| * 500,000 | 2.000 × 10−5 | * 300,000,000 | 3333 × 10−8 |
| * 700,000 | 1.429 × 10−5 | 400,000,000 | 2500 × 10−8 |
| * 900,000 | 1.111 × 10−5 | * 500,000,000 | 2000 × 10−8 |
| 1,200,000 | 8.333 × 10−6 | * 700,000,000 | 1429 × 10−8 |
| * 1,600,000 | 6.250 × 10−6 | * 900,000,000 | 1111 × 10−8 |
| * 2,200,000 | 4.545 × 10−6 | 1,200,000,000 | 8333 × 10−9 |
| * 3,000,000 | 3.333 × 10−6 |
| Tinit | Tfin | Nt | Δt |
|---|---|---|---|
| 0.1 | 0.2 | 500 | 2.000 × 10−4 |
| Nx | Ny | Nx * Ny | Δx * Δy | Nx | Ny | Nx * Ny | Δx * Δy |
|---|---|---|---|---|---|---|---|
| 65 | 65 | 4225 | 2.441 × 10−4 | * 2830 | 2830 | 8,008,900 | 1.249 × 10−7 |
| 100 | 100 | 10,000 | 1.020 × 10−4 | * 3360 | 3360 | 11,289,600 | 8.863 × 10−8 |
| 200 | 200 | 40,000 | 2.525 × 10−5 | 4000 | 4000 | 16,000,000 | 6.253 × 10−8 |
| * 238 | 238 | 56,644 | 1.780 × 10−5 | * 4520 | 4520 | 20,430,400 | 4.897 × 10−8 |
| * 283 | 283 | 80,089 | 1.257 × 10−5 | * 5100 | 5100 | 26,010,000 | 3.846 × 10−8 |
| * 336 | 336 | 112,896 | 8.911 × 10−6 | * 5760 | 5760 | 33,177,600 | 3.015 × 10−8 |
| 400 | 400 | 160,000 | 6.281 × 10−6 | 6500 | 6500 | 42,250,000 | 2.368 × 10−8 |
| * 452 | 452 | 204,304 | 4.916 × 10−6 | * 7240 | 7240 | 52,417,600 | 1.908 × 10−8 |
| * 510 | 510 | 260,100 | 3.860 × 10−6 | * 8060 | 8060 | 64,963,600 | 1.540 × 10−8 |
| * 576 | 576 | 331,776 | 3.025 × 10−6 | * 8980 | 8980 | 80,640,400 | 1.240 × 10−8 |
| 650 | 650 | 422,500 | 2.374 × 10−6 | 10,000 | 10,000 | 100,000,000 | 1.000 × 10−8 |
| * 724 | 724 | 524,176 | 1.913 × 10−6 | * 11,900 | 11,900 | 142,000,000 | 7.074 × 10−9 |
| * 806 | 806 | 649,636 | 1.543 × 10−6 | * 14,100 | 14,100 | 199,000,000 | 5.037 × 10−9 |
| * 898 | 898 | 806,404 | 1.243 × 10−6 | * 16,800 | 16,800 | 282,000,000 | 3.547 × 10−9 |
| 1000 | 1000 | 1,000,000 | 1.002 × 10−6 | 20,000 | 20,000 | 400,000,000 | 2.503 × 10−9 |
| * 1190 | 1190 | 1,416,100 | 7.074 × 10−7 | * 23,800 | 23,800 | 566,440,000 | 1.767 × 10−9 |
| * 1410 | 1410 | 1,988,100 | 5.037 × 10−7 | * 28,300 | 28,300 | 800,890,000 | 1.249 × 10−9 |
| * 1680 | 1680 | 2,822,400 | 3.547 × 10−7 | * 33,600 | 33,600 | 112,896,000 | 8.858 × 10−10 |
| 2000 | 2000 | 4,000,000 | 2.503 × 10−7 | 40,000 | 40,000 | 1,600,000,000 | 6.250 × 10−10 |
| * 2380 | 2380 | 5,664,400 | 1.767 × 10−7 |
| Nt | Δt | Nt | Δt |
|---|---|---|---|
| 1 | 1.000 × 100 | 5000 | 2.000 × 10−4 |
| 5 | 2.000 × 10−1 | 50,000 | 2.000 × 10−5 |
| 50 | 2.000 × 10−2 | * 200,000 | 5.000 × 10−6 |
| 500 | 2.000 × 10−3 |
| Nx | Δx | Nx | Δx |
|---|---|---|---|
| 40 | 2564 × 10−1 | 120,000 | 8333 × 10−5 |
| 120 | 8403 × 10−2 | 400,000 | 2500 × 10−5 |
| 400 | 2506 × 10−2 | * 1,200,000 | 8333 × 10−6 |
| 1200 | 8340 × 10−3 | * 4,000,000 | 2500 × 10−6 |
| 4000 | 2501 × 10−3 | ** 12,000,000 | 8333 × 10−7 |
| 12,000 | 8334 × 10−4 | ** 40,000,000 | 2500 × 10−7 |
| 40,000 | 2500 × 10−4 |
| Nt | Δt | Nt | Δt |
|---|---|---|---|
| 1 | 1.000 × 10−1 | 5000 | 2.000 × 10−5 |
| 5 | 2.000 × 10−2 | 50,000 | 2.000 × 10−6 |
| 50 | 2.000 × 10−3 | * 200,000 | 5.000 × 10−7 |
| 500 | 2.000 × 10−4 |
| Nx | Ny | Nx * Ny | Δx * Δy | Nx | Ny | Nx * Ny | Δx * Δy |
|---|---|---|---|---|---|---|---|
| 65 | 65 | 4225 | 2441 × 10−4 | * 1000 | 1000 | 1,000,000 | 1002 × 10−6 |
| 100 | 100 | 10,000 | 1020 × 10−4 | * 2000 | 2000 | 4,000,000 | 2503 × 10−7 |
| 200 | 200 | 40,000 | 2525 × 10−5 | ** 4000 | 4000 | 16,000,000 | 6253 × 10−8 |
| 400 | 400 | 160,000 | 6281 × 10−6 | ** 6500 | 6500 | 42,250,000 | 2368 × 10−8 |
| 650 | 650 | 422,500 | 2374 × 10−6 |
| Nx | Ny | Nx × Ny | Δx × Δy |
|---|---|---|---|
| 100 | 100 | 10,000 | 1.020 × 10−5 |
| 180 | 180 | 32,400 | 3.121 × 10−5 |
| 320 | 320 | 102,400 | 9.827 × 10−6 |
| 560 | 560 | 313,600 | 3.200 × 10−6 |
| 1000 | 1000 | 1,000,000 | 1.002 × 10−6 |
| Nt | Δt | Nt | Δt |
|---|---|---|---|
| * 1 | 1.000 × 10−3 | 9000 | 1.111 × 10−7 |
| * 2 | 5.000 × 10−4 | 15,000 | 6.667 × 10−8 |
| * 3 | 3.333 × 10−4 | 30,000 | 3.333 × 10−8 |
| * 5 | 2.000 × 10−4 | 50,000 | 2.000 × 10−8 |
| 9 | 1.111 × 10−4 | 90,000 | 1.111 × 10−8 |
| 15 | 6.667 × 10−5 | 150,000 | 6.667 × 10−9 |
| 30 | 3.333 × 10−5 | 300,000 | 3.333 × 10−9 |
| 50 | 2.000 × 10−5 | 500,000 | 2.000 × 10−9 |
| 90 | 1.111 × 10−5 | **** 650,000 | 1.539 × 10−9 |
| 150 | 6.667 × 10−6 | ** 670,000 | 6.667 × 10−9 |
| 300 | 6.667 × 10−6 | ** 690,000 | 3.333 × 10−9 |
| 500 | 2.000 × 10−6 | ** 700,000 | 1.429 × 10−9 |
| 900 | 1.111 × 10−6 | *** 800,000 | 1.250 × 10−9 |
| 1500 | 6.667 × 10−7 | *** 900,000 | 1.111 × 10−9 |
| 3000 | 3.333 × 10−7 | *** 1,000,000 | 1.000 × 10−9 |
| 5000 | 2.000 × 10−7 |
| Spatial Grid | Nr | Δx2 | Last Value of ErrMax |
|---|---|---|---|
| 100 × 100 | 10,000 | 1.020 × 10−4 | 3.083 × 10−5 |
| 180 × 180 | 32,400 | 3.121 × 10−5 | 9.444 × 10−6 |
| 320 × 320 | 102,400 | 9.827 × 10−6 | 2.984 × 10−6 |
| 560 × 560 | 313,600 | 3.200 × 10−6 | 9.922 × 10−7 |
| 1000 × 1000 | 1,000,000 | 1.002 × 10−6 | <unknown> |
| k | R2 | RMSE |
|---|---|---|
| 3.022 × 10−1 | 0.9999982243 | 1.577 × 10−8 |
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content. |
© 2026 by the authors. Licensee MDPI, Basel, Switzerland. This article is an open access article distributed under the terms and conditions of the Creative Commons Attribution (CC BY) license.
Share and Cite
Koics, D.; Kovács, E.; Hornyák, O. GPU-Parallelization of the Numerical Solution of Large-Scale Transient Partial Differential Equations. Computers 2026, 15, 673. https://doi.org/10.3390/computers15100673
Koics D, Kovács E, Hornyák O. GPU-Parallelization of the Numerical Solution of Large-Scale Transient Partial Differential Equations. Computers. 2026; 15(10):673. https://doi.org/10.3390/computers15100673
Chicago/Turabian StyleKoics, Dániel, Endre Kovács, and Olivér Hornyák. 2026. "GPU-Parallelization of the Numerical Solution of Large-Scale Transient Partial Differential Equations" Computers 15, no. 10: 673. https://doi.org/10.3390/computers15100673
APA StyleKoics, D., Kovács, E., & Hornyák, O. (2026). GPU-Parallelization of the Numerical Solution of Large-Scale Transient Partial Differential Equations. Computers, 15(10), 673. https://doi.org/10.3390/computers15100673

