Semantic Data Integration Techniques for Heterogeneous Data Sources

  • Authors

    • Paul Baran Research Scientist, RAND Corporation, United States. Author

    DOI:

    https://doi.org/10.67228/30715717/IJDEIC-2024PII2Y4R

    Published 08-08-2024

  • Semantic Data Integration, Ontology Alignment, Heterogeneous Data Sources, RDF, Owl, Schema Mapping, Data Interoperability, Semantic Web, Knowledge Representation, Data Fusion

    Issue

    Section

    Articles

    How to Cite

    [1]
    P. Baran, “Semantic Data Integration Techniques for Heterogeneous Data Sources”, IJDEIC, vol. 7, no. 2, pp. 01–15, Aug. 2024, doi: 10.67228/30715717/IJDEIC-2024PII2Y4R.
  • Abstract

    Semantic data integration has become essential in data-driven environments where organizations handle large volumes of heterogeneous data from diverse sources such as databases, web services, XML, sensors, and social media. Traditional integration techniques like schema matching and data warehousing are insufficient to address structural, syntactic, and semantic differences across distributed systems. To overcome these challenges, semantic approaches based on ontologies, knowledge representation, and reasoning have gained importance. This work focuses on ontology-based integration, schema alignment, and semantic mediation to enable interoperability. It also examines integration models such as Global-as-View (GAV), Local-as-View (LAV), and hybrid approaches, along with semantic web technologies like RDF, OWL, and SPARQL. A systematic integration methodology is proposed, involving preprocessing, semantic annotation, ontology matching, conflict resolution, and query transformation. Mathematical models are used to define similarity measures and mapping confidence. Experimental evaluation on benchmark datasets shows improved accuracy, precision, recall, data consistency, and query efficiency compared to traditional methods. Despite these advancements, challenges such as scalability, ontology development, and semantic ambiguity remain open research issues.

  • References

    [1] T. Gruber, “A translation approach to portable ontology specifications,” Knowledge Acquisition, vol. 5, no. 2, pp. 199–220, 1993.

    [2] N. F. Noy, “Semantic integration: A survey of ontology-based approaches,” SIGMOD Record, vol. 33, no. 4, pp. 65–70, 2004.

    [3] H. Wache et al., “Ontology-based integration of information—A survey of existing approaches,” in IJCAI Workshop on Ontologies and Information Sharing, 2001, pp. 108–117.

    [4] E. Rahm and P. A. Bernstein, “A survey of approaches to automatic schema matching,” The VLDB Journal, vol. 10, no. 4, pp. 334–350, 2001.

    [5] S. Melnik, H. Garcia-Molina, and E. Rahm, “Similarity flooding: A versatile graph matching algorithm,” in Proc. IEEE ICDE, 2002, pp. 117–128.

    [6] J. Euzenat and P. Shvaiko, Ontology Matching, 2nd ed. Berlin, Germany: Springer, 2013.

    [7] G. Klyne and J. J. Carroll, “Resource Description Framework (RDF): Concepts and abstract syntax,” W3C Recommendation, 2004.

    [8] D. L. McGuinness and F. van Harmelen, “OWL Web Ontology Language overview,” W3C Recommendation, 2004.

    [9] E. Prud’hommeaux and A. Seaborne, “SPARQL query language for RDF,” W3C Recommendation, 2008.

    [10] C. Bizer, T. Heath, and T. Berners-Lee, “Linked data—the story so far,” International Journal on Semantic Web and Information Systems, vol. 5, no. 3, pp. 1–22, 2009.

    [11] A. Doan, A. Halevy, and Z. Ives, Principles of Data Integration. San Francisco, CA, USA: Morgan Kaufmann, 2012.

    [12] F. Naumann and M. Herschel, An Introduction to Duplicate Detection. Morgan & Claypool, 2010.

    [13] X. L. Dong and F. Naumann, “Data fusion: Resolving data conflicts for integration,” Proceedings of the VLDB Endowment, vol. 2, no. 2, pp. 1654–1655, 2009.

    [14] P. Shvaiko and J. Euzenat, “A survey of schema-based matching approaches,” Journal on Data Semantics, vol. 4, pp. 146–171, 2005.

    [15] L. Bleiholder and F. Naumann, “Data fusion,” ACM Computing Surveys, vol. 41, no. 1, pp. 1–41, 2008.

    [16] Gajula, S. (2023). A review of anomaly identification in finance frauds using machine learning system. International Journal of Current Engineering and Technology, 13(6), 568–575. https://ijcet.evegenis.org/index.php/ijcet/article/view/820

  • Downloads