2.1. Basic Assumptions
Regarding the optimization problem of the three-echelon supply chain network structure studied in this paper, the following assumptions are proposed:
(1) The research area is located within mainland China, where the Earth’s curvature is relatively small, and thus it can be considered a plane. The great circle distance between two points calculated by the Haversine formula ignores the influence of terrain fluctuations, road networks, and detours in reality. The distance calculation is primarily used for relative comparison rather than absolute precision.
(2) The supply capacity of all suppliers and the order demand of factories are deterministic and known, and uncertainties such as seasonal fluctuations, random disturbances, unexpected events, and force majeure are not considered during the study period.
(3) Transportation cost is represented as a linear function of route distance and cargo weight or volume, with cargo-specific coefficients. This assumption reflects the information resolution of the competition data: the organizer provides standardized conversion coefficients but does not provide shipment-level truck utilization, contract tiers, full-truckload/less-than-truckload breakpoints, or carrier-specific minimum charges. Introducing unobserved quantity discounts would therefore add parameters that cannot be calibrated consistently. The assumption should be interpreted as an aggregate planning approximation for comparing network configurations over the same study horizon, not as a detailed freight quotation model. In contrast, distribution center operating cost is piecewise because the supplied cost schedule associates throughput bands with step changes in facility resources, labor, and handling capacity. Thus, linear transportation cost and nonlinear center operating cost describe two different cost-generating mechanisms and are not internally inconsistent. A future operational model could replace the linear freight term with lane- and mode-specific tariffs when FTL/LTL and discount data become available.
(4) All goods from each supplier must be and can only be assigned to one distribution center. Splitting a single supplier’s goods for shipment to different distribution centers is not allowed.
2.3. Sets, Parameters, and Variables
The logistics network distribution center location optimization problem is defined on a directed multigraph . Here, represents the set of nodes, consisting of the supplier set S, the candidate distribution center set C, and the factory set F. To reduce model complexity, suppliers that are geographically close can be considered as the same supplier node; similarly, geographically close factories can be considered as the same factory node. represents the set of arcs, including arcs from suppliers to distribution centers and arcs from distribution centers to factories . However, it is noteworthy that, in real-world supply chain transportation networks, there may be multiple alternative paths connecting two points. In this model, we focus on selecting Pareto-optimal arcs, i.e., for a given path, it is impossible to improve the transportation cost of one node without increasing the transportation cost of at least one other node while satisfying all demands, ensuring efficiency and economic viability of the logistics supply chain network and avoiding resource waste under all constraints.
All transportation tasks in the supply chain network are represented by the transportation demand set , where each transportation demand represents the demand of a specific factory for a specific product. Transportation demands are linked to suppliers through product mapping. Therefore, let the product set be , and each product corresponds to a unique supplier S. Also, let represent the demand of factory for product .
In the model, for any two nodes in the established supply chain network, the distance between them is set as , calculated using the formula. In the constructed supply chain network, we only focus on two types of distances: the distance from supplier to distribution center , denoted as , and the distance from distribution center to factory , denoted as .
In the transportation network established by the model, the total transportation volume is generated by all demands. Let
represent the total shipment volume from supplier
, which equals the sum of the demands from all factories requiring products from that supplier:
According to the linear transportation cost assumption, transportation cost is linearly related to distance. Transporting goods the same distance incurs the same cost. In the model, distance is used as the primary calculation unit for transportation cost. Considering the throughput operation cost and construction/operation cost of distribution centers, we define as the fixed construction cost of a distribution center and as the operation cost of a distribution center, which is determined by a piecewise function , where is the throughput of the distribution center.
Define the following decision variables:
Distribution center throughput decision variable
represents the total cargo volume (units) processed by distribution center
. Transportation variable
represents the shipment volume (units) from distribution center
to factory
. Distance violation variable
represents the excess of the actual distance from supplier i to distribution center j over the allowed maximum distance:
Coverage variable represents the proportion of transportation volume satisfying the distance constraint relative to the total transportation volume.
Based on the above, the symbols and variable definitions for this model are shown in
Table 2.
2.4. Objective Function and Constraints
The objective function (5) indicates minimizing the total cost, which consists of three parts: total transportation cost, total operation cost, and total penalty cost. The total transportation cost includes two parts: the first-stage transportation cost from suppliers to distribution centers and the second-stage transportation cost from distribution centers to factories. Specific calculation formulas are given in Equations (7) and (8). The relationship between the transportation cost coefficients and and the cargo type is shown in Equations (30) and (31). The total operation cost is calculated as in Equation (9), where is the distribution center operation cost function, with its specific piecewise form given in Equation (10). The throughput of a distribution center is defined as the sum of supplies from all suppliers assigned to that center, calculated as in Equation (15). The penalty cost is calculated as in Equation (11), and Equation (33) defines the value of .
Equations (12)–(22) are the constraint equations. Equation (12) ensures each supplier must be assigned to exactly one distribution center. Equation (13) states that only selected distribution centers can receive suppliers. To ensure flow balance, we define the flow balance constraint (14), requiring the total cargo received by each distribution center to equal the total cargo it ships to all factories. Equations (16)–(18) are linearized definitions for the distance violation indicator variable . The actual coverage rate is defined as the proportion of cargo transportation volume satisfying the distance constraint to the total transportation volume, calculated by Equation (19). Equation (20) calculates the total demand . Equation (21) is the requirement for the actual coverage rate target. Equations (22) and (23) are non-negativity constraints for continuous decision variables. Equations (24)–(27) are constraints on the types of decision variables. Equations (28) and (29) are distance symmetry and triangle inequality constraints, where the distance between two points is calculated by the formula in Equation (32).