Journal
Integrating high-dimensional censored data under privacy constraints via localized computations
Limited sample size and censoring inherently limit the statistical efficiency of high-dimensional data analysis. While integrating data from multiple sources can enhance estimation efficiency, concerns remain regarding data privacy breaches and between-site heterogeneity. In this paper, we propose a privacy-preserving approach to integrate the high-dimensional right-censored data with source-level heterogeneity. The proposed method is based on the local computation strategy: each site can obtain an integrative estimation based on its local full dataset and the summary statistics from other sites. For each party, this strategy not only meets the data privacy constraints but also maximizes its local data's utilization. Moreover, we introduce a refined procedure for practical use to avoid the shrinkage of the local covariate effect that is unique across all sites. Theoretical results of the proposed estimates including consistency, asymptotic normality and efficiency gains are attained. Simulation experiments demonstrate its superiority over the integrative methods relying solely on summary statistics and the local estimations. The application to multi-source clinical data of ovarian cancer further verifies its practical effectiveness.
A class of semiparametric models for bivariate survival data
We propose a new class of bivariate survival models based on the family of Archimedean copulas with margins modeled by the Yang and Prentice (YP) model. The Ali-Mikhail-Haq (AMH), Clayton, Frank, Gumbel-Hougaard (GH), and Joe copulas are employed to accommodate the dependency among marginal distributions. Baseline distributions are modeled semiparametrically by the Piecewise Exponential (PE) distribution and the Bernstein polynomials (BP). Inference procedures for the proposed class of models are based on the maximum likelihood (ML) approach. The new class of models possesses some attractive features: i) the ability to take into account survival data with crossing survival curves; ii) the inclusion of the well-known proportional hazards (PH) and proportional odds (PO) models as particular cases; iii) greater flexibility provided by the semiparametric modeling of the marginal baseline distributions; iv) the availability of closed-form expressions for the likelihood functions, leading to more straightforward inferential procedures. The properties of the proposed class are numerically investigated through an extensive simulation study. Finally, we demonstrate the versatility of our new class of models through the analysis of survival data involving patients diagnosed with ovarian cancer.
Estimating treatment effects on duration with disease: a principal stratification framework
In clinical research, estimating the average treatment effect is a common goal. However, when treatment effects vary substantially across individuals, it is often more informative to evaluate the treatment effect within subgroups. This paper focuses on causal inference for a duration outcome in a principal stratum-defined as the subgroup of individuals who would experience a positive duration under one treatment. Motivated by the Danish Vulva Cancer Recurrence Study (DaVulvaRec), which compares intensive versus standard follow-up in women treated for vulvar cancer, we examine the effect of intensive follow-up on the time with a cancer recurrence diagnosis. The principal stratum is in this example women who would be diagnosed with cancer recurrence under the intensive follow-up. We present a framework for identifying and estimating the average treatment effect in the principal stratum under a monotonicity assumption and introduce a sensitivity parameter to evaluate the impact of potential violations of this assumption. Using a multi-state model with pseudo-observations, we account for censoring and demonstrate that this approach offers greater statistical power than conventional comparisons between treatment groups. We illustrate the methodology to sample size calculation, the final analysis of the DaVulvaRec study using a simulated data set and an application to data from a randomized study on colon cancer.
Springer Science and Business Media LLC
1380-7870