Next Article in Journal
uavews 0.1.0-Rehearsal: A Synthetic Multisource, Multimodal Spatiotemporal Dataset and Executable Validation Pipeline for Small-UAV Early Warning
Previous Article in Journal
The Relative Age Effect Beyond Geography: Evidence of a Shared Pattern in Senior International Football
 
 
Font Type:
Arial Georgia Verdana
Font Size:
Aa Aa Aa
Line Spacing:
Column Width:
Background:
Data Descriptor

UDT: Unemployment Duration Tables for a Large City in Poland (2007–2024)

by
Beata Bieszk-Stolorz
1,* and
Joanna Olbryś
2,*
1
Institute of Economics and Finance, University of Szczecin, 71-101 Szczecin, Poland
2
Faculty of Computer Science, Bialystok University of Technology, 15-351 Bialystok, Poland
*
Authors to whom correspondence should be addressed.
Data 2026, 11(9), 246; https://doi.org/10.3390/data11090246 (registering DOI)
Submission received: 5 August 2026 / Revised: 12 September 2026 / Accepted: 16 September 2026 / Published: 19 September 2026
(This article belongs to the Section Information Systems and Data Management)

Abstract

The duration of unemployment, along with the unemployment rate, is one of the most important indicators of the labour market. It provides information on the scale of unemployment, along with its nature and socio-economic impacts. This article contributes through a unique database that contains information on the duration of unemployment, i.e., the time from the moment of registration to the moment of de-registration of unemployed people from the labour office. The database is extensive, contains information on a total of 299,668 people, and covers the years 2007–2024. These people were grouped into cohorts according to the year of de-registration. The unit of time analysed is the month. The numerical UDT database (Microsoft Excel) was created on the basis of information from the Poviat Labour Office in Szczecin, Poland. These data allow for the use of survival analysis methods and the inclusion of censored data in the study. Based on these data and the methodology of construction of life tables, unemployment duration tables were prepared, and a quartile analysis of the time of registration with the labour office was performed. The conclusions of such research may be a guide for policymakers to create employment policies and to promote activities supporting the unemployed.
Dataset: The UDT dataset was submitted as Supplementary Materials to this manuscript.
Dataset License: CC BY-SA

1. Summary

In Poland, a registered unemployed person is a person who is unemployed and not engaged in other gainful employment, registered with a poviat labour office, capable and ready to take up full-time employment or other gainful employment, and meets the conditions specified in Article 2, Section 1 of the Act on the Labour Market and Employment Services. These conditions include, among others, being 18 years of age, not having reached retirement age (60 for women, 65 for men), not receiving certain benefits (e.g., permanent social assistance benefits, disability pensions, retirement pensions), and not being subject to any conditions that would exclude unemployment status (e.g., full-time education, arrest, imprisonment, military service). In the case of a disabled person, the ability and willingness to work is limited to at least half-time [1].
The duration of unemployment, in addition to the unemployment rate, is a key indicator of the quality of the functioning of the labour market and the scale of the problem of long-term unemployment. The unemployment rate alone does not show the full picture of the labour market. Two economies may have similar unemployment rates. However, if in one of them unemployed people quickly find a job, and in the other they remain unemployed for many months, then the socio-economic situation of these economies will be significantly different. The duration of unemployment can be studied using methods derived from demographic research—methods of survival analysis. However, the use of these methods is often impossible due to the lack of access to data characterized by certain specifics. This article presents a database of unemployed persons de-registered from the Poviat Labour Office in Szczecin (Poland) in the years 2007–2024. The UDT (Unemployment Duration Tables) database contains information that allows the use of duration analysis methods to study the time of job search. The data were collected in the form of tables of duration in registered unemployment. These tables were constructed by using the methodology of constructing life tables. This made it possible to estimate the function of duration in unemployment for cohorts of people de-registered in the years 2007–2024. The estimated median duration makes it possible to compare the probability of remaining in registered unemployment in subsequent years.
Survival analysis is a methodology used in many different studies in which we are interested in the time of occurrence of events. The name of these methods reflects their origin in demographic studies on human life expectancy. Subsequently, these methods were used in medicine and in reliability theory to study the uptime of machines and equipment. By events, we mean events in the lives of individuals (not only people) that are the subject of scientific research in the fields of medicine, demography, biology, sociology, economics, technical sciences, banking, capital markets, real estate markets, etc. [2,3,4,5,6,7,8,9]. In relation to human life, these can include death, myocardial infarction, marriage, divorce, birth of a child, graduation, or diagnosis of illness. These methods can be applied to anything that changes, develops, or is destroyed. We can study the time of failure-free operation of a device (light bulb, car), the duration of unemployment, the duration of companies, the time of sale of real estate, or the time to pay back a loan. For all such events, it is possible to look for their cause and determine the risk factor. Survival analysis is used as a tool in issues such as the following [10]:
  • Analysis of the effectiveness of disease treatments;
  • Isolation of risk factors to prevent diseases;
  • Evaluation of the reliability of technical equipment;
  • Understanding the mechanisms of biological phenomena;
  • Monitoring social phenomena such as divorce and unemployment.
In labour market research, survival analysis methods can be applied to the analysis of economic activity [11], the duration of unemployment [12,13,14], and the time of employment [15]. There have been publications on the use of life tables in unemployment surveys. However, as these studies did not consider the uniform distribution of censorship over time intervals, they were limited to a shortened follow-up period and concerned only one selected year [16,17]. Existing, publicly available datasets contain aggregated data that prevents the use of the survival analysis methods. To fill this gap, we present a set of actual data containing information on the duration of registered unemployment. These data descriptors detail the methodology, data structure, and potential applications for researchers and decision-makers monitoring the labour market. The described methodology could theoretically be adapted to other socio-economic phenomena in which it is important to study the duration of a given phenomenon. Due to the large number of years analysed, including the financial crisis in 2007–2008 and the period of the COVID–19 pandemic in 2020–2022, these data can help researchers better understand the impact of crisis situations on the labour market.

2. Problem Formulation

This section presents the methodological background concerning survival analyses. The survival function and median duration, as well parametric and non-parametric models, are described in detail.

2.1. Survival Function

The basis of survival analysis is the duration, which is considered as a non-negative random variable T. The probability distribution for time T is described by five basic functions: the cumulative density function, the probability density function, the survival function, the hazard function, and the cumulative hazard function [18,19]. The primary function in survival analysis is the survival function. It determines the probability that the duration for a given unit will be longer than t. It can be written as follows:
S ( t ) = P ( T > t ) = 1 F ( t ) ,
where
  • T—Duration;
  • F ( t ) —Cumulative distribution function of random variable T.
It is a non-increasing, non-negative function, taking values from 0 to 1. It satisfies the conditions S ( 0 ) = 1 and lim t + S ( t ) = 0 .
In the case of discontinuity, it is right-continuous. In other words, the survival function determines the probability that an interesting event will not occur until t. The survival curve, which is a graph of survival function, gives the expected percentage of individuals for whom the event has not yet occurred by time t [10]. This curve starts at 100% of the study population and shows the percentage of the population that continues to survive in subsequent periods for as long as information is available. It can be used not only for survival as such but also for maintaining a state free of disease, complications, or another endpoint [20].
In practice, using real data, we usually obtain graphs that are step functions. Because the study period is never infinite and there may be competing risks of failure, it is possible that not all subjects will experience the event. The estimated survival function may therefore not fall completely to zero at the end of the study [21].
Due to the diversity of analysed phenomena, it is worth remembering that the event ending the observation does not always have to be negative in its essence. In studies on human life expectancy or in failure-free analysis, it is desirable that the survival time is as long as possible. The event that ends the observation is the death of a person or a device failure. However, in the case of the duration of unemployment or the time to sell real estate, a shorter duration is desirable, and the event ending the observation (taking up a job or selling real estate) has a positive significance for the analysed phenomenon. However, in order to standardize terminology, there we can often talk about the risk of an event occurring in general.

2.2. Median Duration

Since the distribution of survival times is characterized by a high skewness, location parameters of the structure are preferred. Once the survival function has been estimated, it is easy to obtain an estimate of the median survival time. This is the time M, after which the chance of survival is equal to 0.5, or S ( M ) = 0.5 . Since non-parametric estimates of survival function S ( t ) are the stepped ones, it is usually not possible to obtain an estimated survival time that will make the survival function exactly equal to 0.5. In this case, the median survival is defined as the smallest observed survival time for which the value of the estimated survival function is less than 0.5. It can be written as follows [22]:
M = m i n { t i | S ( t i ) < 0.5 } ,
where t i is the observed survival time of i-th person, i = 1 , 2 , n .
In a particular case, for which the estimated survival function is exactly equal to 0.5 for all values of t belonging to a certain interval from t j to t j + 1 , the median can be estimated as the midpoint of this interval, or M = t j + t j + 1 2 . If there is no censored survival time, the median survival time will be the smallest time after which the probability of survival is equal to 0.5. In the case of censored observations, it happens that the estimated survival function is greater than 0.5 for all values of t. In these cases, the median survival cannot be estimated. A procedure similar to the one described above can be used to estimate other percentiles of the survival time distribution [22,23].

2.3. Parametric and Non-Parametric Models

Among the methods used in survival analysis, parametric, semi-parametric, and non-parametric models are distinguished. Parametric survival models are those in which the duration of a phenomenon (survival time presented by a random variable T) is assumed to follow a known distribution. The construction of these models requires the theoretical distribution of the analysed variable. This method is used in the case of human life expectancy analysis. This is because distributions such as exponential, Weibull, Pareto, Gompertz, and logarithmic-normal have been known for a long time and describe human life expectancy well [24,25,26]. In socio-economic phenomena, the distribution of survival time is usually unknown. This is the main reason for using non-parametric and semi-parametric methods. The oldest way of determining a non-parametric survival function estimator is to construct life tables, and the most popular non-parametric survival function estimator is the Kaplan–Meier estimator [27]. Among semi-parametric methods, Cox’s proportional hazard model [28] is used to model hazards. Non-parametric approaches are better in the sense that they are more flexible and can avoid errors in model specification. However, parametric models have the advantage that the parameters can be interpreted [29].

3. Data Description

The data and their application in the analysis of duration in unemployment are closely related to the specificity of survival analysis methods. The advantage of these methods is the possibility of using censored data in research. Their presence is important because of their impact on an individual’s likelihood of survival until censorship.

3.1. Censored Data

Classical survival analysis focuses on the time to the occurrence of a single event for each individual analysed. Specifically, it is the time that has elapsed from the initiating event to the event that completes the observation (final event). This time is called survival time, even if the endpoint is something other than death. This is a fundamental problem that we almost always encounter when using survival analysis methods. In fact, we have to wait for an event to occur. When the research ends, it usually turns out that the event in question occurred for some of the individuals studied but did not occur for others. The data obtained in this way are a mixture of complete and incomplete observations. Incomplete observations are called the censored survival times. In this case, the censored survival times are called right-censored [30]. The set of individuals (people) for whom an event did not occur before a given moment t and who were not censored before the moment t is called the set of risk at the moment t [10]. It may also be the case that people who originally participated in the study are lost from observation and their survival status (survivor/deceased) cannot be determined. A similar situation is also present when the people in the sample die from causes completely unrelated to the disease under study; these people can also be considered lost from observation. Such observations are also right-censored.
In general, most methods derived from survival analysis take into account a key problem of data, called censorship. Censorship occurs when we have some information about the survival time of a given individual but we do not know the exact survival time. In fact, there can be three types of censorship: left-sided (no information about the start of the process), right-sided (no information about the end of the process), or two-way (no information about both the beginning and the end of the process). In economic research, we usually deal with right-sided censorship, because the observation ends earlier than the moment of occurrence of the event, which ends the process for a group of individuals [31]. It can be assumed that the survival time is right-censored at time t if it is known that it is greater than t [32]. Not including censored observations in the survival analysis becomes a potential source of error. Until censorship, i.e., the loss of information to an individual, it is part of the risk set and affects the probability of occurrence of an event.

3.2. Life Tables

It is assumed that the life table is probably the earliest statistical tool used to study human mortality. Edmund Halley (1693) and John Graunt (1662) independently developed the first life expectancy tables based on the populations in Wrocław (now in Poland) and England. This is a well-organized description of mortality rates by age. The life table method (also known as the actuarial method or the Cutler–Ederer method) is an approximation of the Kaplan–Meier method [33]. It is based on grouped survival times and is suitable for big data. Modern methods used in survival analysis have reduced the research significance of life tables. However, despite this, they are fundamental to understanding survival data [30]. Using the traditional life table method for a specific period, a non-parametric estimation of the following can be obtained:
  • Survival function;
  • Cumulative density function;
  • Probability density function;
  • Hazard function.
There are two types of life tables: cohort and period life tables. Period life tables show mortality rates in a given period for a specific population. A cohort life table is built on data collected by registering the survival time from the birth of the first member of the population to the death of the last member. A cohort is a term used in statistics and other sciences (e.g., demography, medicine), denoting a set of objects—most often people—isolated from a population due to an event or process occurring simultaneously for the entire set. A cohort should be distinguished based on statistically significant features and homogeneous in terms of them. A cohort is a group of people with common, defining characteristics (typically people who have experienced a shared event over a selected period, such as the birth of a child or graduation). Studies using cohorts are called cohort studies. Estimating survival function based on a life table is also known as actuarial estimation of survival function. It is obtained by first dividing the observation period into several time intervals. These intervals do not have to be of equal length [22].
These people are observed over time, and their event time or censoring time is recorded so that it falls on one of k + 1 contiguous, non-overlapping intervals: I 1 = [ t 0 , t ] for j = 1 and I j = ( t j 1 , t j ] for j = 2 , , k , k + 1 . We assume that t 0 = 0 and, in particular, t k + 1 = .
The basic construction of the table is described below:
  • First Column: t j 1 —Beginning of interval I j .
  • Second Column: t j —End of interval I j .
  • Third Column: n j —Number of people in the interval I j .
  • Fourth Column: c j —Number of censored observations in the interval I j .
  • Fifth Column: n j —An estimate of the number of individuals at risk of experiencing the event in the interval I j . We assume that the censored survival times occur evenly throughout the j-th interval, so that the average number of people at risk in this interval is as follows [34]:
    n j = n j c j 2 .
  • Sixth Column: d j —Number of people who experienced an event in the interval I j .
  • Seventh Column: S ( t j 1 ) —Value of the survival function in the interval I j . It is assumed that
    S ( t 0 ) = S ( 0 ) = 1 ,
    S ( t j ) = S ( t j 1 ) · ( 1 d j n j . ) for j = 1 , 2 , , k .
A graphical estimate of the survival function will be a step function with constant values of the function in each time interval. According to the designations adopted above, the life table can be in the following form [22,35,36] (Table 1):
The form of the estimated survival function obtained by this method is sensitive to the selection of intervals used in its construction. On the other hand, life expectancy estimation is particularly useful in situations where the actual time of death is unknown and the only information available is the number of deaths and the number of censored observations that occur in a series of successive time intervals. In practice, interval-censored survival data occur quite frequently.
Once the actual survival time is known, one can use the estimate from the life table. However, grouping survival times causes some loss of information. In this case, alternative non-parametric methods for estimating survival functions, such as the Kaplan–Meier estimator [27], are more suitable. In demography, the life table method for estimating survival functions continues to be at the forefront for describing human mortality. A lifespan-based method or an actuarial method may be better for large datasets or when the timing of events is inaccurate [37].
Application of this method is not limited to human populations. It can also be applied to any life form, as well as to inanimate objects [38,39,40,41,42].

4. Methods

This section describes the methodology of construction of the numerical UDT database (Microsoft Excel). The UDT file contains cohort unemployment duration tables for Szczecin, a large city in Poland, within the years 2007–2024.

4.1. Data Used in the Research

The database comes from the Poviat Labour Office in Szczecin. Szczecin is a city in northwestern Poland, located about 13–15 km from the Polish–German border. It is a port city with a developed shipbuilding industry. Szczecin has 385.2 thousand inhabitants (2025) and has an area of 301 km2. Individual data collected by poviat labour offices in Poland is not published. It is only stored in labour offices. Currently (since 2011), all poviat labour offices in Poland use the SyriuszStd IT system. There is no central unit that collects such data from all offices in Poland. Only aggregated data is sent to the Statistical Office in Poland. Information on the number of people de-registered in individual months, the registered unemployment rate, and various statistical studies can be found on the websites of labour offices. However, their publication depends on the activity of the given office. More detailed data for scientific research purposes can be obtained from the office upon request.
The database contains data on unemployed persons de-registered from the office in the years 2007–2024. The initial total number of observations was 302,313. The data received from the employment office needed to be cleaned up. The dataset contained some errors that had been corrected before the tables were created. In the survival analysis, it is assumed that the duration of the phenomenon for the unit t > 0 . In the data received, there were observations for which t = 0 . This was probably due to errors made by officials or a failure to complete certain formalities at the time of registration. From the output data, 2645 records with a duration of zero were removed. Finally, 299,668 records were used to create the tables. The number of deleted records (per year) is included in Table A1 (Appendix A).
In this way, 18 cohorts of de-registered people were obtained. The dataset is a descriptive dataset of completed unemployment spells grouped by year of exit. However, the creation of such cohorts is of great importance to labour offices. Labour offices are interested in information about their activities in a given year. They want to know whether they managed to ‘clear’ the queue of long-term unemployed people in that year, and whether those taking up employment are doing so quickly. The data in each cohort includes information about the period of registration (in months) and the reason for de-registration. The reasons for de-registration are different and have been divided into two groups: taking up a job, and de-registration for other reasons. Taking up a job indicates taking up employment with an employer, starting a business, or taking up subsidized work. De-registration for other reasons includes resignation from co-operation with the office (one of the most common reasons), retirement, starting education, starting military service, going to another city, or going abroad. The numbers of individual cohorts, the share of unemployed people taking up employment, and the maximum time of registration are presented in Table 2. In addition, Table 2 includes a column containing the registered unemployment rates in Szczecin. Their comparison with the obtained results allows us to understand the reasons for the occurrence of outliers.
The obtained data allow for the use of survival analysis methods and the construction of cohort unemployment duration tables.

4.2. Cohort Unemployment Duration Tables

In this case, the random variable T is the time from the moment of registration to the moment of de-registration of the unemployed person from the labour office. Thus, it describes the duration of registered unemployment. These are not consecutive calendar months of the year but subsequent months of stay in the register. The study was conducted with the use of survival analysis to assess the chance of an unemployed person taking up employment (or the risk of remaining in the register) by the expiry of the registration period. Incomplete observations are right-censored. One month was taken as a unit of time. The first time interval (the first month), according to the adopted designations, is I 1 = [ 0 , 1 ] , and j-th time interval (j-th month) is I j = ( j 1 , j ] . The number of months in each cohort was different (see Table 2).
The described life table construction methodology makes it possible to construct the unemployment duration tables. In the case of the random variable T specified in this way, the basic quantities necessary to construct the unemployment duration cohort tables can be defined as follows:
  • Complete observation is the time from the moment of registration of an unemployed person to the moment of de-registration due to taking up employment.
  • Censored observation is the time from the moment an unemployed person is registered to the moment of de-registration for a reason other than taking up employment.
  • Description of the columns in the unemployment duration table:
    • First Column (1): t j 1 —Beginning of the j-th month.
    • Second Column (2): t j —End of the j-th month.
    • Third Column (3): n j —Number of unemployed persons at the moment j 1 .
    • Fourth Column (4): c j —Number of people de-registered for reasons other than taking up employment in the j-th month.
    • Fifth Column (5): n j —Average number of people with a chance to take up employment in the j-th month.
    • Sixth Column (6): d j —Number of people who took up employment in the j-th month.
    • Seventh Column (7): S ( t j 1 ) —Value of survival function in the j-th month.
The data provided by the labour office included the following variables: k, n j , and d j . Using Formulae (3)–(5), it was possible to determine the remaining elements of the unemployment duration tables. Based on the determined survival function and Formulae (2), (6) and (7), it was possible to calculate the median and, similarly, the quartiles of the unemployment duration. For the values n j , c j , and d j defined in this way, the survival function describes the probability of remaining in the cohort, i.e., in the register of unemployed persons. Hence, it is correct to call such a survival function the function of unemployment duration, and the life table can be called the unemployment duration table. The assumption of censorship is also used in medical science. A person who has died of causes other than the one being studied is considered to be censored because of the disease in question [43]. The function of duration in unemployment determined on the basis of tables is stepped and non-increasing. On this basis, quartiles of duration in unemployment can be determined. For the analysed data, all quartiles for all years exist; they are shown in Figure 1.
Quartiles allow us to compare the speed of exiting the unemployment register in individual years. On their basis, the following become possible:
  • Assessment of the variation in the duration of unemployment based on the value of the interquartile range (IQR):
    I Q R = Q 3 Q 1 .
    The higher the IQR, the greater the dispersion of the time of being unemployed.
  • Evaluation of distribution skewness based on the following:
    • Comparison of values of lower-quartile spread M Q 1 and the upper-quartile spread Q 3 M ,
    • Bowley’s quartile skewness coefficient:
      A Q = ( Q 3 M ) ( M Q 1 ) Q 3 Q 1 .
  • Identifying long-term unemployment on the basis of the third quartile.
  • Comparisons between groups within a single cohort or between cohorts.
In the analysis of the duration of unemployment, it is sometimes more informative to compare M Q 1 and Q 3 M or to calculate Bowley’s coefficient than to give the IQR alone. Table 3 presents the values of interquartile range, lower-quartile spread, upper-quartile spread, and Bowley’s quartile skewness coefficients for time of registration of unemployed persons with the labour office between 2007 and 2024.

4.3. Baseline Results

A quartile analysis of the data presented in Table 3 and the registered unemployment rate from Table 2 leads to the following conclusions:
  • The high IQR for the 2007 and 2008 cohorts indicates a high variation in the duration of unemployed people belonging to these cohorts. In these years, it is much higher than for other cohorts.
  • In all analysed cohorts, Q 3 M > M Q 1 , which indicates a clear positive skew in the distribution of the unemployed duration. The distributions are therefore positively skewed, which is often the case with unemployment duration. It follows that most people leave the unemployment register relatively quickly, but some remain in it as registered unemployed for a very long time.
  • The positive skew of the distributions is confirmed by Bowley’s quartile skewness coefficient, which is positive in all analysed cohorts.
  • The high value of third quartiles in 2007–2008 means that a significant proportion of the long-term unemployed persons took up employment in these years. In 2007–2008, people de-registered at the labour office had a longer median duration than in the case of other cohorts. These are the years of a decrease in the registered unemployment rate in Szczecin, from 11.8% in 2006 to 6.5% in 2007 and 4.3% in 2008. A longer time to work among people de-registered in these years testifies to the effect of “unloading the queue”. People who had had trouble finding a job in previous years finally found one.
  • The low value of quartiles in 2020 is noteworthy. This is the first year of the pandemic. The short time to register to work in the 2020 cohort is well explained by the specificity of the labour market in Szczecin, which is a large border city where a large proportion of people work in Germany, mainly in services. These include people with lower qualifications (salespeople, construction workers, caregivers), but also people with high qualifications (doctors, pharmacists, nurses, teachers, and engineers of various specialities). These are very often people who live in Szczecin and commute to work every day. The closure of borders due to the pandemic made it impossible to perform work, as not every type of work could be performed remotely. Such people often registered with the office and took up employment in Szczecin.
When using a dataset, its limitations must be taken into consideration. The data are of a regional nature, as they were obtained from one Polish city with county rights (the Szczecin city county). However, there is an advantage in this limitation: It is regional data that are most needed for economic analyses, particularly analyses of the labour market in a given country. They allow one to isolate areas with worse or better conditions, and they enable effective creation of social policy. The data can be used for comparisons with other regions in Poland or with other similar cities abroad. The data presented in the form of tables are suitable to determine other quantities as part of the survival analysis.
Another limitation was the lack of detailed information on other reasons for de-registration, which could have been treated as competing events. It was also not possible to construct cohorts of people registered in specific years. The labour office provided data only on the year of de-registration. This data did not include the registration history of individual people, due to a concern for complete anonymity.

5. User Notes

The UDT dataset is provided as Supplementary Materials in Excel format. The file is divided into separate sheets for each year from 2007 to 2024. Each sheet contains seven columns, as described in Section 4.2.
Data in these sheets allow the construction of additional components of the unemployment duration tables based on grouped observations. These components include the following [40,44]:
  • Estimated standard errors of the survival function;
  • Estimated probability of taking up employment during the analysed period, given unemployment at its start;
  • Estimated probability of not taking up employment during the analysed period, given unemployment at its start;
  • Estimated hazard and probability density functions, and the corresponding estimated standard errors of estimates;
  • Estimated probability of taking up employment within a given number of months after registration;
  • Statistical comparison of unemployment duration tables using non-parametric tests.

Supplementary Materials

The following supporting information can be downloaded at https://www.mdpi.com/article/10.3390/data11090246/s1. The Supplementary Materials contain the UDT dataset in Excel format. The file is divided into separate sheets for each year from 2007 to 2024. The dataset contains the unemployment duration tables for Szczecin (Poland).

Author Contributions

Conceptualization, methodology, B.B.-S.; software, B.B.-S. and J.O.; validation, B.B.-S. and J.O.; formal analysis, B.B.-S. and J.O.; investigation, B.B.-S. and J.O.; resources, B.B.-S.; data curation, B.B.-S.; writing—original draft preparation, B.B.-S. and J.O.; writing—review and editing, B.B.-S. and J.O.; visualization, B.B.-S. and J.O.; supervision, B.B.-S. and J.O.; funding acquisition, B.B.-S. and J.O. All authors have read and agreed to the published version of the manuscript.

Funding

The research contribution of the second named author was supported by project no. WZ/WI-IIT/2/25 at Bialystok University of Technology and financed from the research subsidy of Bialystok University of Technology.

Institutional Review Board Statement

Not applicable.

Informed Consent Statement

Not applicable.

Data Availability Statement

The UDT dataset was submitted as Supplementary Materials to this manuscript.

Acknowledgments

The study was carried out as a part of a research internship at the Bialystok University of Technology entitled “The use of entropy in the research of socio-economic phenomena”, in collaboration between the University of Szczecin (Poland) and the Bialystok University of Technology (Poland) (15 June–31 December 2026).

Conflicts of Interest

The authors declare no conflicts of interest.

Abbreviations

The following abbreviations are used in this manuscript:
UDTUnemployment Duration Tables
IQRInterquartile range
AQBowley’s quartile skewness coefficient

Appendix A. The UDT—Additional Information

Table A1 reports the additional information about the erroneous records in the dataset received from the employment office (per year). In particular, the second column presents the total number of de-registered persons before removing errors, while the last column documents the total number of de-registered persons after removing the errors (these numbers are also included in Table 2).
Table A1. The number of people de-registered before and after the deletion of erroneous data.
Table A1. The number of people de-registered before and after the deletion of erroneous data.
YearTotal Number of De-Registered Persons Before Removing ErrorsNumber of Erroneous RecordsFraction of Erroneous DataTotal Number of De-Registered Persons After Removing Errors
200723,8861410.59%23,745
200817,23640.02%17,232
200919,39800.00%19,398
201017,7931801.01%17,613
201115,211170.11%15,194
201215,614440.28%15,570
201323,9712090.87%23,762
201424,7232801.13%24,443
201525,8813131.21%25,568
201623,7342871.21%23,447
201719,9312351.18%19,696
201815,0972241.48%14,873
201912,8761961.52%12,680
202079001281.62%7772
202189801021.14%8878
20229629920.96%9537
20239358991.06%9259
202411,095940.85%11,001
Total302,31326450.87%299,668

References

  1. Sejm, R.P. Ustawa z Dnia 20 Marca 2025 r. o Rynku Pracy i służBach Zatrudnienia (Dziennik Ustaw 2025, poz. 620). 2025. Available online: https://eli.gov.pl/api/acts/DU/2025/620/text/T/D20250620L.pdf (accessed on 28 August 2026).
  2. Meyer, B.D. Unemployment insurance and unemployment spells. Econometrica 1990, 58, 757–782. [Google Scholar] [CrossRef] [Scilit]
  3. Box-Steffensmeier, J.M.; Zorn, C.J.W. Duration models and proportional hazards in political science. Am. J. Political Sci. 2001, 45, 972–988. [Google Scholar] [CrossRef] [Scilit]
  4. Roszkowska, P.; Prorokowski, L. Model of financial crisis contagion: A survey-based simulation by means of the modified Kaplan-Meier survival plots. Folia Oecon. Stetin. 2013, 13, 22–55. [Google Scholar] [CrossRef] [Scilit]
  5. Sączewska-Piotrowska, A. Poverty duration of households of the self-employed. Econom. Ekon. Adv. Appl. Data Anal. 2015, 1, 44–55. [Google Scholar] [CrossRef] [Scilit]
  6. Markowicz, I. Duration analysis of firms–cohort tables and hazard function. Int. J. Bus. Soc. Res. 2015, 5, 36–47. [Google Scholar]
  7. Bieszk-Stolorz, B.; Dmytrów, K. A survival analysis in the assessment of the influence of the SARS-CoV-2 pandemic on the probability and intensity of decline in the value of stock indices. Eurasian Econ. Rev. 2021, 11, 363–379. [Google Scholar] [CrossRef] [Scilit]
  8. Putek-Szeląg, E.; Gdakowicz, A. Application of duration analysis methods in the study of the exit of a real estate sale offer from the offer database system. In Data Analysis and Classification. Methods and Applications; Jajuga, K., Najman, K., Walesiak, M., Eds.; Springer: Cham, Switzerland, 2021; pp. 153–169. [Google Scholar] [CrossRef] [Scilit]
  9. Wycinka, E. Competing risk models of default in the presence of early repayments. Econom. Ekon. Adv. Appl. Data Anal. 2019, 23, 99–120. [Google Scholar] [CrossRef] [Scilit]
  10. Aalen, O.O.; Borgan, O.; Gjessing, H.K. Survival and Event History Analysis: A Process Point of View; Springer: New York, NY, USA, 2008. [Google Scholar]
  11. Landmesser, J. The survey of economic activity of people in rural areas-the analysis using the econometric hazard models. Acta Univ. Lodz. Folia Oecon. 2009, 228, 385–392. [Google Scholar]
  12. Landmesser, J. Econometric analysis of unemployment duration using hazard models. Stud. Ekon. 2009, 1–2, 79–92. [Google Scholar]
  13. Güell, M.; Lafuente, C. Revisiting the determinants of unemployment duration: Variance decomposition à la ABS in Spain. Labour Econ. 2022, 78, 102233. [Google Scholar] [CrossRef] [Scilit]
  14. Grzenda, W. Estimating the probability of leaving unemployment for older people in Poland using survival models with censored data. Stat. Transit. New Ser. 2023, 24, 241–256. [Google Scholar] [CrossRef] [Scilit]
  15. Basha, L.; Gjika, E. Accelerated failure time models in analyzing duration of employment. J. Phys. Conf. Ser. 2022, 2287, 012014. [Google Scholar] [CrossRef] [Scilit]
  16. Bieszk-Stolorz, B.; Markowicz, I. Unemployment duration tables. Acta Univ. Lodz. Folia Oecon. 2016, 5, 77–85. [Google Scholar] [CrossRef] [Scilit]
  17. Bieszk-Stolorz, B.; Markowicz, I. Variants of exiting nemployment–duration tables. In Sustainable Economic Development and Advancing Education Excellence in the Era of Global Pandemic: Proceedings of the 36th International Business Information Management Association, IBIMA, Granada, Spain, 4–5 November 2020; Soliman, K.S., Ed.; International Business Information Management Association: Norristown, PA, USA, 2020; pp. 1466–1475. [Google Scholar]
  18. Kalbfleisch, J.D.; Prentice, R.L. The Statistical Analysis of Failure Time Data; John Wiley & Sons, Inc.: New York, NY, USA, 2011. [Google Scholar]
  19. Cameron, A.C.; Trivedi, P.K. Microeconometrics: Methods and Applications; Cambridge University Press: New York, NY, USA, 2005. [Google Scholar]
  20. Dicker, R.C.; Coronado, F.; Koo, D.; Parrish, R.G. Principles of Epidemiology in Public Health Practice: An Introduction to Applied Epidemiology and Biostatistics, 3rd ed.; Centers for Disease Control and Prevention: Atlanta, GA, USA, 2006.
  21. Kleinbaum, D.G.; Klein, M. Survival Analysis: A Self-Learning Text, 3rd ed.; Springer: New York, NY, USA, 2012. [Google Scholar]
  22. Collett, D. Modelling Survival Data in Medical Research, 4th ed.; Chapman and Hall/CRC: New York, NY, USA, 2023. [Google Scholar]
  23. Machin, D.; Cheung, Y.B.; Parmar, M. Survival Analysis: A Practical Approach; John Wiley & Sons, Inc.: New York, NY, USA, 2006. [Google Scholar]
  24. Gompertz, B. On the nature of the function expressive of the law of human mortality, and on a new mode of determining the value of life contingencies. Philos. Trans. R. Soc. Lond. 1825, 115, 513–583. [Google Scholar] [CrossRef] [Scilit]
  25. Missov, T.I.; Vaupel, J.W. Mortality implications of mortality plateaus. SIAM Rev. 2015, 57, 61–70. [Google Scholar] [CrossRef] [Scilit]
  26. Zhang, Z. Parametric regression model for survival data: Weibull regression model as an example. Ann. Transl. Med. 2016, 4, 484. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  27. Kaplan, E.L.; Meier, P. Nonparametric estimation from incomplete observations. J. Am. Stat. Assoc. 1958, 53, 457–481. [Google Scholar] [CrossRef] [Scilit]
  28. Cox, D.R. Regression models and life tables (with discussion). J. R. Stat. Soc. B 1972, 34, 187–220. [Google Scholar]
  29. Läuter, H.; Liero, H. Nonparametric estimation and testing in survival models. In Probability, Statistics and Modelling in Public Health; Nikulin, M.S., Commenges, D., Huber, C., Eds.; Springer: Boston, MA, USA, 2006; pp. 319–331. [Google Scholar]
  30. Selvin, S. Survival Analysis for Epidemiologic and Medical Research; Cambridge University Press: New York, NY, USA, 2008. [Google Scholar]
  31. Cappelli, C.; Zhang, H. Survival trees. In Statistical Methods for Biostatistics and Related Fields; Härdle, W., Mori, Y., Vieu, P., Eds.; Springer: Berlin/Heidelberg, Germany, 2007; pp. 167–179. [Google Scholar]
  32. Vittinghoff, E.; Glidden, D.V.; Shiboski, S.C.; McCulloch, C.E. Regression Methods in Biostatistics: Linear, Logistic, Survival, and Repeated Measures Models; Springer: New York, NY, USA, 2004. [Google Scholar]
  33. Stevenson, M. An Introduction to Survival Analysis; EpiCentre, IVABS, Massey University: Palmerston North, New Zealand, 2007. [Google Scholar]
  34. Hosmer, D.W., Jr.; Lemeshow, S.; May, S. Applied Survival Analysis: Regression Modeling of Time-to-Event Data; John Wiley & Sons, Inc.: New York, NY, USA, 2008. [Google Scholar]
  35. Gehan, E.A. Estimating survival functions from the life table. J. Chronic Dis. 1969, 21, 629–644. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  36. Klein, J.P.; Moeschberger, M.L. Survival Analysis: Techniques for Censored and Truncated Data, 2nd ed.; Springer: New York, NY, USA, 2003. [Google Scholar]
  37. Allison, P.D. Survival Analysis Using SAS: A Practical Guide; SAS Institute: Cary, NC, USA, 2010. [Google Scholar]
  38. Sakagami, S.F.; Fukuda, H. Life tables for worker honeybees. Popul. Ecol. 1968, 10, 127–139. [Google Scholar] [CrossRef] [Scilit]
  39. Spinage, C.A. African ungulate life tables. Ecology 1972, 53, 645–652. [Google Scholar] [CrossRef] [Scilit]
  40. Namboodiri, K.; Suchindran, C.M. Life Table Techniques and Their Applications; Academic Press, Inc.: Orlando, FL, USA, 1987. [Google Scholar]
  41. Schoen, R.; Nelson, V.E. Marriage, divorce, and mortality: A life table analysis. Demography 1974, 11, 267–290. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  42. Sullivan, D.F. A single index of mortality and morbidity. HSMHA Health Rep. 1971, 86, 347. [Google Scholar] [CrossRef] [Scilit]
  43. Zelen, M. Forward and backward recurrence times and length biased sampling: Age specific models. Lifetime Data Anal. 2004, 10, 325–334. [Google Scholar] [CrossRef] [Scilit] [PubMed]
  44. Lawless, J.F. Statistical Models and Methods for Lifetime Data, 2nd ed.; John Wiley & Sons, Inc.: New York, NY, USA, 2003. [Google Scholar]
Figure 1. Unemployment duration quartiles determined on the basis of registered unemployment duration tables in Szczecin in the years 2007–2024.
Figure 1. Unemployment duration quartiles determined on the basis of registered unemployment duration tables in Szczecin in the years 2007–2024.
Data 11 00246 g001
Table 1. Life table.
Table 1. Life table.
t j 1 t j n j c j n j d j S ( t j 1 )
t 0 t 1 n 1 c 1 n 1 d 1 S ( t 0 )
t 1 t 2 n 2 c 2 n 2 d 2 S ( t 1 )
t k t k + 1 n k c k n k d k S ( t k )
Table 2. Total number of de-registered persons, fraction of de-registered persons to work, maximum time to de-registration, and registered unemployment rate in Szczecin in 2007–2024.
Table 2. Total number of de-registered persons, fraction of de-registered persons to work, maximum time to de-registration, and registered unemployment rate in Szczecin in 2007–2024.
YearTotal Number of De-Registered PersonsFraction of De-Registered Persons to WorkMaximum Time to De-Registration (Months)Registered Unemployment Rate in Szczecin
200723,74534%1616.50%
200817,23232%1654.30%
200919,39836%1868.50%
201017,61341%1929.70%
201115,19439%1679.90%
201215,57039%19011.00%
201323,76246%19010.60%
201424,44345%1759.30%
201525,56843%1806.80%
201623,44742%2174.70%
201719,69640%2213.10%
201814,87342%2482.60%
201912,68044%2452.40%
2020777269%1963.90%
2021887863%2493.30%
2022953751%2123.10%
2023925958%1603.60%
202411,00153%1973.40%
Source: Own elaboration on the basis of data from the Poviat Labour Office in Szczecin.
Table 3. Interquartile range, lower-quartile spread, upper-quartile spread, and Bowley’s quartile skewness coefficient.
Table 3. Interquartile range, lower-quartile spread, upper-quartile spread, and Bowley’s quartile skewness coefficient.
YearIQRM-Q1Q3-MAQ
20079733640.32
200810421830.60
2009305250.67
2010144100.43
2011217140.33
2012268180.38
2013309210.40
2014369270.50
2015429330.57
2016577500.75
2017615560.84
2018354310.77
2019203170.70
20207250.43
2021155100.33
2022256190.52
2023224180.64
2024274230.70
Disclaimer/Publisher’s Note: The statements, opinions and data contained in all publications are solely those of the individual author(s) and contributor(s) and not of MDPI and/or the editor(s). MDPI and/or the editor(s) disclaim responsibility for any injury to people or property resulting from any ideas, methods, instructions or products referred to in the content.

Share and Cite

MDPI and ACS Style

Bieszk-Stolorz, B.; Olbryś, J. UDT: Unemployment Duration Tables for a Large City in Poland (2007–2024). Data 2026, 11, 246. https://doi.org/10.3390/data11090246

AMA Style

Bieszk-Stolorz B, Olbryś J. UDT: Unemployment Duration Tables for a Large City in Poland (2007–2024). Data. 2026; 11(9):246. https://doi.org/10.3390/data11090246

Chicago/Turabian Style

Bieszk-Stolorz, Beata, and Joanna Olbryś. 2026. "UDT: Unemployment Duration Tables for a Large City in Poland (2007–2024)" Data 11, no. 9: 246. https://doi.org/10.3390/data11090246

APA Style

Bieszk-Stolorz, B., & Olbryś, J. (2026). UDT: Unemployment Duration Tables for a Large City in Poland (2007–2024). Data, 11(9), 246. https://doi.org/10.3390/data11090246

Article Metrics

Back to TopTop