Graph-Based Approaches to Complex Data Transformation Workflows

  • Authors

    • Tony Hoare University of Oslo, Norway Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2022PII4H3A

    Published 10-06-2022

  • Graph-based modeling, Data transformation workflows, Workflow optimization, DAGs (Directed Acyclic Graphs), Data pipelines, ETL/ELT systems, Data engineering, Workflow orchestration, Graph traversal

    Issue

    Section

    Articles

    How to Cite

    [1]
    T. Hoare, “Graph-Based Approaches to Complex Data Transformation Workflows”, IJDEIC, vol. 5, no. 2, pp. 01–09, Oct. 2022, doi: 10.67228/30715717/IJDEIC-2022PII4H3A.
  • Abstract

    Modern data workflows are increasingly complex, involving heterogeneous data sources, dynamic transformation logic, and stringent performance requirements. Traditional linear and script-based approaches often struggle to provide the modularity, scalability, and traceability required for these workflows. In this paper, we explore graph-based approaches to modeling and executing complex data transformation workflows. By representing transformation logic as graphs where nodes denote operations and edges represent data or control dependencies these methods offer a flexible and powerful paradigm for structuring, optimizing, and managing workflows. We present a comprehensive framework for graph-based workflow modeling, discuss optimization strategies, and examine real-world applications and systems that leverage this methodology. We also analyze the trade-offs and limitations of such approaches and suggest future research directions in dynamic and intelligent graph-based systems.

  • References

    [1] Abadi, D. J., Marcus, A., Madden, S. R., & Hollenbach, K. J. (2009). Scalable semantic web data management using vertical partitioning. Proceedings of the VLDB Endowment, 1(1), 411–422.

    [2] Armbrust, M., et al. (2015). Spark SQL: Relational data processing in Spark. Proceedings of the 2015 ACM SIGMOD International Conference on Management of Data, 1383–1394.

    [3] Binnig, C., Kossmann, D., & Kraska, T. (2016). How is the weather tomorrow?: Towards a benchmark for the cloud. CIDR.

    [4] Chen, Z., & Akella, A. (2017). Scheduling and resource allocation in DAG-based workflow systems. Journal of Parallel and Distributed Computing, 110, 1–14.

    [5] Chikhi, R., & Medini, T. (2017). Towards lineage tracking in graph-based ETL workflows. Data Engineering Workshops (ICDEW), 112–119.

    [6] Garlan, D., & Shaw, M. (1994). An introduction to software architecture. In Advances in Software Engineering and Knowledge Engineering (Vol. 1).

    [7] Kandel, S., et al. (2011). Wrangler: Interactive visual specification of data transformation scripts. Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, 3363–3372.

    [8] Katsiapis, A., Sattler, K.-U., & Nüske, N. (2020). Optimization of graph workflows with parallel execution. Information Systems, 92, 101525.

    [9] Krishnan, K., et al. (2014). Data lineage: What, why, and how? Proceedings of the 2014 ACM SIGMOD International Conference on Management of Data, 5–12.

    [10] Li, Y., & Parthasarathy, S. (2018). Workflow provenance: Modeling, querying, and visualization. Journal of Data and Information Quality, 10(2), 7:1–7:26.

    [11] Malewicz, G., et al. (2010). Pregel: A system for large-scale graph processing. Proceedings of the 2010 ACM SIGMOD International Conference on Management of Data, 135–146.

    [12] Ramesh, M., & Narayan, A. (2021). AI-driven data transformation discovery using graph models. IEEE Transactions on Knowledge and Data Engineering.

    [13] Zaharia, M., et al. (2016). Apache Spark: A unified engine for big data processing. Communications of the ACM, 59(11), 56–65.

  • Downloads