dc.creatorMontezanti,Diego
dc.creatorFrati,Fernando Emmanuel
dc.creatorRexachs,Dolores
dc.creatorLuque,Emilio
dc.creatorNaiouf,Marcelo
dc.creatorDe Giusti,Armando
dc.date2012-12-01
dc.date.accessioned2023-09-25T18:35:14Z
dc.date.available2023-09-25T18:35:14Z
dc.identifierhttp://www.scielo.edu.uy/scielo.php?script=sci_arttext&pid=S0717-50002012000300006
dc.identifier.urihttps://repositorioslatinoamericanos.uchile.cl/handle/2250/8838438
dc.descriptionThe challenge of improving the performance of current processors is achieved by increasing the integration scale. This carries a growing vulnerability to transient faults, which increase their impact on multicore clusters running large scientific parallel applications. The requirement for enhancing the reliability of these systems, coupled with the high cost of rerunning the application from the beginning, create the motivation for having specific software strategies for the target systems. This paper introduces SMCV, which is a fully distributed technique that provides fault detection for message-passing parallel applications, by validating the contents of the messages to be sent, preventing the transmission of errors to other processes and leveraging the intrinsic hardware redundancy of the multicore. SMCV achieves a wide robustness against transient faults with a reduced overhead, and accomplishes a trade-off between moderate detection latency and low additional workload.
dc.formattext/html
dc.languageen
dc.publisherCentro Latinoamericano de Estudios en Informática
dc.rightsinfo:eu-repo/semantics/openAccess
dc.sourceCLEI Electronic Journal v.15 n.3 2012
dc.subjecttransient fault
dc.subjectsilent data corruption
dc.subjectmulticore cluster
dc.subjectparallel scientific application
dc.subjectsoft error detection
dc.subjectmessage content validation
dc.subjectreliability
dc.titleSMCV: a Methodology for Detecting Transient Faults in Multicore Clusters
dc.typeinfo:eu-repo/semantics/article


Este ítem pertenece a la siguiente institución