Fault-Tolerant Data Transformation Mechanisms in Heterogeneous Data Environments

  • Authors

    • Dr. K. Balasubramanian Professor, Bharathiar University, India. Author

    DOI:

    https://doi.org/10.67228/30713315/IJAIDT-2020PI7D4T

    Published 03-05-2020

  • Fault Tolerance, Data Transformation, Heterogeneous Data Environments, Data Quality, Fault Detection, Adaptive Pipelines, Data Provenance, Data Integration, Robust Data Processing

    Issue

    Section

    Articles

    How to Cite

    [1]
    B. K, “Fault-Tolerant Data Transformation Mechanisms in Heterogeneous Data Environments”, IJAIDT, vol. 3, no. 1, pp. 01–13, Mar. 2020, doi: 10.67228/30713315/IJAIDT-2020PI7D4T.
  • Abstract

    In heterogeneous data environments, where data originates from diverse sources and formats, ensuring reliable and accurate data transformation is critical yet challenging. Faults during data transformation processes can lead to corrupted outputs, reduced data quality, and compromised decision-making. This paper presents fault-tolerant data transformation mechanisms designed specifically to address the complexities introduced by heterogeneous data sources. We propose a comprehensive framework incorporating advanced fault detection, adaptive processing pipelines, and recovery techniques that enhance the robustness and reliability of data transformation workflows. Through experimental evaluation on varied datasets, our approach demonstrates improved fault resilience, reduced downtime, and higher data integrity compared to existing methods. The results highlight the potential of fault-tolerant mechanisms to significantly enhance data processing reliability in complex, heterogeneous data ecosystems.

  • References

    [1] Stonebraker, M., & Çetintemel, U. (2005). "One size fits all": An idea whose time has come and gone. Proceedings of the 21st International Conference on Data Engineering, 2–11.

    [2] Abadi, D. J., et al. (2003). Aurora: A new model and architecture for data stream management. The VLDB Journal, 12(2), 120-139.

    [3] Wu, E., & Rundensteiner, E. A. (2003). Fault-tolerant stream processing using replicated state machines. Proceedings of the 19th International Conference on Data Engineering, 444-455.

    [4] Chen, Y., et al. (2019). Adaptive data transformation pipelines in heterogeneous environments. IEEE Transactions on Knowledge and Data Engineering, 31(7), 1309-1322.

    [5] Hellerstein, J. M., et al. (2007). The datacenter as a computer: An introduction to the design of warehouse-scale machines. Synthesis Lectures on Computer Architecture, 4(1), 1-108.

    [6] Bu, Y., et al. (2009). Naiad: A timely dataflow system. Proceedings of the Twenty-Fourth ACM Symposium on Operating Systems Principles, 439-455.

    [7] Simmhan, Y., Plale, B., & Gannon, D. (2005). A survey of data provenance in e-science. ACM SIGMOD Record, 34(3), 31-36.

    [8] Grolinger, K., et al. (2013). Data management in cloud environments: NoSQL and NewSQL data stores. Journal of Cloud Computing, 2(1), 22.

    [9] Fernandez, A., et al. (2014). A survey on fault tolerance in cloud computing. Journal of Network and Computer Applications, 51, 1-17.

    [10] Jagadish, H. V., et al. (2014). Making database systems usable. Foundations and Trends® in Databases, 6(1-2), 1-141.

    [11] Kotsifakos, V., et al. (2020). Adaptive fault tolerance in stream processing engines. IEEE Transactions on Parallel and Distributed Systems, 31(12), 2889-2903.

    [12] Venkataraman, S., et al. (2016). Apache flink: Stream and batch processing in a single engine. Proceedings of the VLDB Endowment, 9(13), 1717-1720.

    [13] Chen, J., et al. (2018). Data quality management in big data systems: A systematic literature review. Information Systems, 74, 47-70.

    [14] Stonebraker, M., et al. (2018). The architecture of a database system. Foundations and Trends® in Databases, 1(2), 141-259.

  • Downloads