Open Access
December 2023 Optimal subgroup selection
Henry W. J. Reeve, Timothy I. Cannings, Richard J. Samworth
Author Affiliations +
Ann. Statist. 51(6): 2342-2365 (December 2023). DOI: 10.1214/23-AOS2328


In clinical trials and other applications, we often see regions of the feature space that appear to exhibit interesting behaviour, but it is unclear whether these observed phenomena are reflected at the population level. Focusing on a regression setting, we consider the subgroup selection challenge of identifying a region of the feature space on which the regression function exceeds a pre-determined threshold. We formulate the problem as one of constrained optimisation, where we seek a low-complexity, data-dependent selection set on which, with a guaranteed probability, the regression function is uniformly at least as large as the threshold; subject to this constraint, we would like the region to contain as much mass under the marginal feature distribution as possible. This leads to a natural notion of regret, and our main contribution is to determine the minimax optimal rate for this regret in both the sample size and the Type I error probability. The rate involves a delicate interplay between parameters that control the smoothness of the regression function, as well as exponents that quantify the extent to which the optimal selection set at the population level can be approximated by families of well-behaved subsets. Finally, we expand the scope of our previous results by illustrating how they may be generalised to a treatment and control setting, where interest lies in the heterogeneous treatment effect.

Funding Statement

The second author was supported by Engineering and Physical Sciences Research Council (EPSRC) New Investigator Award EP/V002694/1.
The third author was supported by Engineering and Physical Sciences Research Council (EPSRC) Programme Grant EP/N031938/1, EPSRC Fellowship EP/P031447/1 and European Research Council Advanced Grant 101019498.


We thank the anonymous reviewers for constructive feedback that helped to improve the paper.


Download Citation

Henry W. J. Reeve. Timothy I. Cannings. Richard J. Samworth. "Optimal subgroup selection." Ann. Statist. 51 (6) 2342 - 2365, December 2023.


Received: 1 February 2023; Revised: 1 September 2023; Published: December 2023
First available in Project Euclid: 20 December 2023

MathSciNet: MR4682700
zbMATH: 07783618
Digital Object Identifier: 10.1214/23-AOS2328

Primary: 62G05

Keywords: FWER , nonparametric inference , selective inference , Subgroup selection

Rights: Copyright © 2023 Institute of Mathematical Statistics

Vol.51 • No. 6 • December 2023
Back to Top