A mathematical programming approach for integrated multiple linear regression subset selection and validation
DC Field | Value | Language |
---|---|---|
dc.contributor.author | Chung, S. | - |
dc.contributor.author | Park, Y.W. | - |
dc.contributor.author | Cheong, T. | - |
dc.date.accessioned | 2021-08-31T19:21:03Z | - |
dc.date.available | 2021-08-31T19:21:03Z | - |
dc.date.created | 2021-06-17 | - |
dc.date.issued | 2020 | - |
dc.identifier.issn | 0031-3203 | - |
dc.identifier.uri | https://scholar.korea.ac.kr/handle/2021.sw.korea/60740 | - |
dc.description.abstract | Subset selection for multiple linear regression aims to construct a regression model that minimizes errors by selecting a small number of explanatory variables. Once a model is built, various statistical tests and diagnostics are conducted to validate the model and to determine whether the regression assumptions are met. Most traditional approaches require human decisions at this step. For example, the user may repeat adding or removing a variable until a satisfactory model is obtained. However, this trial-and-error strategy cannot guarantee that a subset that minimizes the errors while satisfying all regression assumptions will be found. In this paper, we propose a fully automated model building procedure for multiple linear regression subset selection that integrates model building and validation based on mathematical programming. The proposed model minimizes mean squared errors while ensuring that the majority of the important regression assumptions are met. We also propose an efficient constraint to approximate the constraint for the coefficient t-test. When no subset satisfies all of the considered regression assumptions, our model provides an alternative subset that satisfies most of these assumptions. Computational results show that our model yields better solutions (i.e., satisfying more regression assumptions) compared to the state-of-the-art benchmark models while maintaining similar explanatory power. © 2020 Elsevier Ltd | - |
dc.language | English | - |
dc.language.iso | en | - |
dc.publisher | Elsevier Ltd | - |
dc.subject | Errors | - |
dc.subject | Mathematical programming | - |
dc.subject | Mean square error | - |
dc.subject | Model buildings | - |
dc.subject | Set theory | - |
dc.subject | Benchmark models | - |
dc.subject | Computational results | - |
dc.subject | Explanatory power | - |
dc.subject | Explanatory variables | - |
dc.subject | Mean squared error | - |
dc.subject | Multiple linear regressions | - |
dc.subject | Satisfactory modeling | - |
dc.subject | Traditional approaches | - |
dc.subject | Linear regression | - |
dc.title | A mathematical programming approach for integrated multiple linear regression subset selection and validation | - |
dc.type | Article | - |
dc.contributor.affiliatedAuthor | Cheong, T. | - |
dc.identifier.doi | 10.1016/j.patcog.2020.107565 | - |
dc.identifier.scopusid | 2-s2.0-85089000582 | - |
dc.identifier.bibliographicCitation | Pattern Recognition, v.108 | - |
dc.relation.isPartOf | Pattern Recognition | - |
dc.citation.title | Pattern Recognition | - |
dc.citation.volume | 108 | - |
dc.type.rims | ART | - |
dc.type.docType | Article | - |
dc.description.journalClass | 1 | - |
dc.description.journalRegisteredClass | scie | - |
dc.description.journalRegisteredClass | scopus | - |
dc.subject.keywordPlus | Errors | - |
dc.subject.keywordPlus | Mathematical programming | - |
dc.subject.keywordPlus | Mean square error | - |
dc.subject.keywordPlus | Model buildings | - |
dc.subject.keywordPlus | Set theory | - |
dc.subject.keywordPlus | Benchmark models | - |
dc.subject.keywordPlus | Computational results | - |
dc.subject.keywordPlus | Explanatory power | - |
dc.subject.keywordPlus | Explanatory variables | - |
dc.subject.keywordPlus | Mean squared error | - |
dc.subject.keywordPlus | Multiple linear regressions | - |
dc.subject.keywordPlus | Satisfactory modeling | - |
dc.subject.keywordPlus | Traditional approaches | - |
dc.subject.keywordPlus | Linear regression | - |
dc.subject.keywordAuthor | Mathematical programming | - |
dc.subject.keywordAuthor | Regression diagnostics | - |
dc.subject.keywordAuthor | Subset selection | - |
Items in ScholarWorks are protected by copyright, with all rights reserved, unless otherwise indicated.
(02841) 서울특별시 성북구 안암로 14502-3290-1114
COPYRIGHT © 2021 Korea University. All Rights Reserved.
Certain data included herein are derived from the © Web of Science of Clarivate Analytics. All rights reserved.
You may not copy or re-distribute this material in whole or in part without the prior written consent of Clarivate Analytics.