CATS: cache-aware task scheduling for Hadoop-based systems
DC Field | Value | Language |
---|---|---|
dc.contributor.author | Lim, Byungnam | - |
dc.contributor.author | Kim, Jong Wook | - |
dc.contributor.author | Chung, Yon Dohn | - |
dc.date.accessioned | 2021-09-02T22:46:33Z | - |
dc.date.available | 2021-09-02T22:46:33Z | - |
dc.date.created | 2021-06-16 | - |
dc.date.issued | 2017-12 | - |
dc.identifier.issn | 1386-7857 | - |
dc.identifier.uri | https://scholar.korea.ac.kr/handle/2021.sw.korea/81433 | - |
dc.description.abstract | Today with the explosion of big data, data-intensive cluster computing systems have driven to a new data processing paradigm. As Hadoop, one of the most famous data processing frameworks, achieves high performance by running multiple tasks in parallel across nodes in large clusters, task scheduling is considered as one of the most important factors affecting the overall performance. In modern operating systems, caching is used to improve local disk access times, providing data from the main memory without disk accesses. This option, however, is poorly utilized by existing task scheduling methods of Hadoop-based systems, mainly due to the inability of tracking cached data in shared-nothing distributed environments. In this paper, we propose a cache-aware task scheduling method, cache-aware task scheduling (CATS), for Hadoop-based systems which is able to exploit the operating system's buffer cache and assign tasks to nodes in consideration of the cached data. Through comprehensive experiments, we show that the proposed cache-aware scheduling improves the overall job execution time for various workload types and data sizes. | - |
dc.language | English | - |
dc.language.iso | en | - |
dc.publisher | SPRINGER | - |
dc.title | CATS: cache-aware task scheduling for Hadoop-based systems | - |
dc.type | Article | - |
dc.contributor.affiliatedAuthor | Chung, Yon Dohn | - |
dc.identifier.doi | 10.1007/s10586-017-0920-6 | - |
dc.identifier.scopusid | 2-s2.0-85019596165 | - |
dc.identifier.wosid | 000414780400071 | - |
dc.identifier.bibliographicCitation | CLUSTER COMPUTING-THE JOURNAL OF NETWORKS SOFTWARE TOOLS AND APPLICATIONS, v.20, no.4, pp.3691 - 3705 | - |
dc.relation.isPartOf | CLUSTER COMPUTING-THE JOURNAL OF NETWORKS SOFTWARE TOOLS AND APPLICATIONS | - |
dc.citation.title | CLUSTER COMPUTING-THE JOURNAL OF NETWORKS SOFTWARE TOOLS AND APPLICATIONS | - |
dc.citation.volume | 20 | - |
dc.citation.number | 4 | - |
dc.citation.startPage | 3691 | - |
dc.citation.endPage | 3705 | - |
dc.type.rims | ART | - |
dc.type.docType | Article | - |
dc.description.journalClass | 1 | - |
dc.description.journalRegisteredClass | scie | - |
dc.description.journalRegisteredClass | scopus | - |
dc.relation.journalResearchArea | Computer Science | - |
dc.relation.journalWebOfScienceCategory | Computer Science, Information Systems | - |
dc.relation.journalWebOfScienceCategory | Computer Science, Theory & Methods | - |
dc.subject.keywordAuthor | Task scheduling | - |
dc.subject.keywordAuthor | Distributed systems | - |
dc.subject.keywordAuthor | Hadoop | - |
dc.subject.keywordAuthor | In-memory | - |
Items in ScholarWorks are protected by copyright, with all rights reserved, unless otherwise indicated.
(02841) 서울특별시 성북구 안암로 14502-3290-1114
COPYRIGHT © 2021 Korea University. All Rights Reserved.
Certain data included herein are derived from the © Web of Science of Clarivate Analytics. All rights reserved.
You may not copy or re-distribute this material in whole or in part without the prior written consent of Clarivate Analytics.