Source-linked AI summary

The Survey of Data Mining Applications And Feature Scope

Neelamadhab Padhy, Dr. Pragnyaban Mishra, Rasmita Panigrahi

arXiv:1211.5723v1cs.DBcs.IR

TL;DR

The paper addresses how organizations can analyze and use massive, heterogeneous datasets for decision-making. It surveys data mining techniques, tasks, life cycles, systems, and applications across domains, concluding that the field offers broad application and research scope. Its healthcare discussion also makes clear that application success depends on clean, standardized, and shareable data.

  • Problem

    Organizations generate massive data in varied formats but need tools to extract useful information, patterns, and knowledge for decision-making.

  • Method

    The paper reviews data mining techniques, tasks, system architectures, life cycles, visualization, methods, applications, and future research directions.

  • Results

    The review identifies broad data mining applications across fields, including healthcare, education, manufacturing, security, sports, taxation, and product quality.

  • Takeaways & Limitations

    Data mining provides a broad framework for extracting patterns and supporting decisions across varied data sources and application domains.

  • Takeaways & Limitations

    Healthcare data mining success depends on clean healthcare data, with further benefits requiring improved capture, storage, preparation, standardization, and sharing.

Abstract

from arXiv · show

In this paper we have focused a variety of techniques, approaches and different areas of the research which are helpful and marked as the important field of data mining Technologies. As we are aware that many Multinational companies and large organizations are operated in different places of the different countries.Each place of operation may generate large volumes of data. Corporate decision makers require access from all such sources and take strategic decisions.The data warehouse is used in the significant business value by improving the effectiveness of managerial decision-making. In an uncertain and highly competitive business environment, the value of strategic information systems such as these are easily recognized however in todays business environment,efficiency or speed is not the only key for competitiveness.This type of huge amount of data are available in the form of tera-topeta-bytes which has drastically changed in the areas of science and engineering.To analyze,manage and make a decision of such type of huge amount of data we need techniques called the data mining which will transforming in many fields.This paper imparts more number of applications of the data mining and also focuses scope of the data mining which will helpful in the further research.

1. INTRODUCTION

The paper frames data mining as a response to overwhelming, heterogeneous data that organizations struggle to convert into useful information for decision-making. It introduces data mining concepts, tasks, methods, life cycles, visualizations, applications, and future research scope.

  • Organizations are data rich but information poor because large collections in formats such as audio, video, text, numbers, and figures are difficult to interpret.The paper emphasizes that retrieval alone is insufficient for extracting useful knowledge and supporting managerial decisions.
  • Data mining extracts hidden predictive information and patterns from large databases to help organizations focus on important information.The paper presents data mining as a tool for automatic summarization, essence extraction, pattern discovery, and decision support.
  • The paper surveys data mining concepts, tasks, classification, life cycles, visualization, methods, applications, and future directions across seven sections.The final section reviews applications and proposes feature directions and scope for further research.

2. The Data Mining Task

The data mining task encompasses exploratory analysis, pattern and rule discovery, content-based retrieval, system classification, and an iterative project life cycle. Tasks and systems vary with data sources, discovered knowledge, techniques, and user interaction.

  • Exploratory Data Analysis: Exploratory data analysis interactively examines large repositories, including when users do not know exactly what they are searching for.It serves both searching without predefined knowledge and analyzing available data.
  • Discovering Patterns and Rules: Pattern and rule discovery identifies hidden patterns in clusters using rule induction and clustering algorithms such as K-Means and K-Medoids.The stated aim is to determine how best to detect patterns across groups of data objects.
  • Retrieval by Content: Content-based retrieval finds frequently used or similar patterns in audio, video, and image datasets.The task searches datasets for patterns resembling a pattern of interest.
  • Types of Data Mining System: Data mining systems are classified by data source, discovered knowledge or functionality, techniques used, and degree of user interaction.Functionalities include characterization, discrimination, association, classification, and clustering, while techniques include machine learning, neural networks, statistics, and visualization.
  • Data Mining Life Cycle: A data mining project follows six non-rigid phases, moving between stages as outcomes require, from business understanding through deployment.The phases include defining objectives, collecting data, modeling, evaluation, and presenting or implementing results.

5. Visualizing Data Mining Model

The paper presents visualization as a way to make hidden data mining models understandable and trustworthy, alongside predictive and descriptive models and a range of mining methods. It also discusses process support for selecting data, algorithms, and valid knowledge discovery workflows.

  • Visualizing Data Mining Model: Data visualization helps users understand and trust models whose information is hidden in repositories.Its stated objective is to provide an overall view of the data mining model.
  • Data Mining Models: Predictive models estimate unknown values from known values, whereas descriptive models identify patterns, relationships, and data properties.Examples include classification and regression for prediction, and clustering, summarization, and association rules for description.
  • New way to define the KDD Process: Knowledge discovery in databases is defined as identifying valid, novel, potentially useful, and ultimately understandable patterns in data.The paper associates KDD with data preparation, pattern search, knowledge judgment, and evaluation.
  • Data Mining Methods: Data mining methods include OLAP, classification, clustering, association rule mining, temporal and time-series analysis, spatial mining, and web mining.These methods use varied data sources and algorithm families, including statistical, decision-tree, nearest-neighbor, neural-network, and genetic algorithms.
  • Data Mining Applications: Knowledge discovery assistants can enumerate valid processes, rank alternatives, and support knowledge sharing during data mining.The paper also describes attempts to automate data and algorithm selection in generalized data mining tools.

7.1 Data Mining Applications in Healthcare

Healthcare data mining has substantial potential, but its success depends on clean, well-prepared healthcare data. The paper identifies standardization, inter-organizational sharing, text mining, and image analysis as directions for expanding applications.

  • Healthcare Data Requirements: Healthcare data mining can support useful applications, but its success hinges on the availability of clean healthcare data.The paper emphasizes improving how healthcare data are captured, stored, prepared, and mined.
  • Healthcare Data Requirements: Standardizing clinical vocabulary and sharing data across organizations are proposed to enhance healthcare data mining benefits.These directions address preparation and interoperability of healthcare information.
  • Expanding Healthcare Data Mining: Text mining can expand healthcare analysis beyond quantitative data such as doctors’ notes and clinical records.Healthcare applications may combine mixed data types before mining text.
  • Expanding Healthcare Data Mining: Image analysis, including MRI scans, is identified as another direction for healthcare data mining.The paper notes that progress has been made in incorporating images into healthcare applications.

7.2 Data mining is used for market basket analysis

Data mining is presented as a tool for discovering customer purchasing associations in market basket analysis and as an emerging trend in education.

  • Market basket analysis uses data mining to find associations among products customers place in their shopping baskets.Retailers can use these purchasing patterns to identify customer buying intentions and promote business.
  • Data mining is identified as an emerging trend in education systems worldwide.
  • Educational data mining can reveal patterns, associations, and anomalies that support decision-making in higher education.The paper associates these uses with improved efficiency, retention, promotion, learning outcomes, and reduced process costs.

7.4 Data mining is now used in many different areas in manufacturing engineering

Manufacturing engineering uses data mining to analyze production data, identify hidden patterns, and support process and product-quality improvements, while current methods remain difficult to apply and interpret.

  • Manufacturing data mining can identify errors, improve design methods, enhance data quality, and support decision-making.
  • Analyzing hidden patterns in manufacturing data can help control processes and enhance product quality.
  • Mining manufacturing data is tedious, and generated knowledge can be difficult to interpret because relationship identification is complex.
  • CRISP-DM is proposed as a methodology providing high-level instructional steps for applying data mining in engineering.
  • Further research is needed to develop generic guidelines for varied data and problem types in manufacturing engineering.

7.5 Data Mining Applications can be generic or domain specific.

Data mining systems may be generic or domain specific: generic systems support method selection and interpretation, while domain-specific applications target particular data and objectives across varied fields.

  • Generic data mining applications guide data selection, method selection, and result interpretation, while multi-agent systems can automate technique selection.
  • A proposed multi-tier data mining system is intended to enhance data mining process performance.
  • Multi-tier systems include user interface, data mining services, data access services, and data, with one-tier, two-tier, and three-tier architectures.
  • CRM data mining research reviews applications and commonly used techniques, but the review does not claim to be exhaustive.
  • Domain-specific applications use targeted data and algorithms to generate knowledge for specific objectives.
  • For foreign-accented French identification, Logistic Regression was found to be the most robust among 20 applied algorithms.
  • Automatic linguistic profiling using lexical and syntactic features achieved 97% accuracy in selecting the correct text author.

7.10 In Medical Science

The paper describes data mining applications in medical science, web education, and credit-risk evaluation, including diagnostic support and analysis of large educational and financial datasets.

  • Medical data mining applications include disease diagnosis, healthcare, patient profiling, and medical-history generation.
  • Neural networks with back-propagation and association-rule mining are used for tumor classification in mammograms.
  • Prediction algorithms achieved 100% accuracy for 91.3% of observed cases in diagnosing lung abnormalities.
  • Web-education data mining discovers relationships in student usage data to guide courseware improvements.Teachers or course authors can use the discovered relationships to decide which modifications may improve effectiveness.
  • Credit-risk evaluation applies data mining to large consumer-credit datasets that are difficult to analyze economically and manually.

7.13 The Intrusion Detection in the Network

The paper presents data mining as useful for analyzing network traffic and identifying anomalies, while also surveying applications across sports and crime analysis.

  • The Intrusion Detection in the Network: Classification methods distinguish normal from abnormal network traffic for intrusion detection.A TCP header outside existing clusters can be treated as an anomaly.
  • Sports data Mining: Data mining tools are applied to maintain and retrieve large volumes of sports data as needed.The section frames sports as an application beyond business use.
  • Sports data Mining: Data mining supports sports scouting, player selection, coaching, training, strategy planning, and squad optimization.A Bayesian classifier is described for predicting Cy Young Award winners.
  • Crime data mining: Crime data mining analyzes large criminal and terrorist-activity databases using clustering, classification, and string comparison.Applications include linking persons, organizations, vehicles, and deceptive information in records.

7.17 The data mining system implemented at the Internal Revenue Service

This section describes data mining applications in tax enforcement, product-quality improvement, e-commerce, digital libraries, and engineering prediction.

  • Tax enforcement: An Internal Revenue Service data mining system identified high-income individuals involved in abusive tax shelters and ranked potentially abusive transactions.The system combined relationship visualization with data mining.
  • Product quality: Regression and neural-network models achieved accuracy above 80% for improving product quality, with the neural network performing better.The passage links these methods to discovering previously unknown rules and reducing cost.
  • E-commerce: E-commerce provides plentiful, reliable data and measurable returns, enabling data mining to support knowledge generation and business decisions.The passage describes integration of e-commerce and data mining as improving results.
  • Digital libraries: Digital libraries are suitable for data mining because they contain diverse electronic resources, including text, images, video, audio, pictures, and maps.Data mining can support finding, collecting, storing, and preserving digital data.
  • Engineering prediction: Data mining produced 100% correct predictions on an engineering test file with nine features.The paper describes applications including engineering cost estimation and design decisions.

8. The Scope of Data Mining

The paper scopes data mining as the extraction of valuable information and predictive patterns from large data stores, using diverse techniques and varying user interaction.

  • Automated prediction of trends and behaviors: Data mining automates predictive analysis in large databases, including targeted marketing and forecasting bankruptcy or default.Past promotional data can identify targets likely to maximize future mailing return on investment.
  • Data mining techniques: The scope includes neural networks, decision trees, evolutionary optimization, and nearest-neighbor classification.These methods respectively support nonlinear prediction, rule-generating decisions, evolutionary search, and similarity-based classification.
  • Scope and purpose: Data mining extracts useful if-then rules and valuable business information from large databases by searching for statistically significant patterns.The paper compares this process with mining valuable ore from immense material.
  • Scope and limitations: The paper concludes that method and data selection require domain knowledge because no completely generic data mining system has been found.Domain experts remain necessary to apply systems and generate required knowledge.

Authors

The authors are information-technology academics and researchers whose interests include data warehousing, data mining, distributed systems, databases, neural networks, algorithms, and cryptography.

  • Authors: Mr. Neelamadhab is an information-technology assistant professor pursuing doctoral research in data mining.His interests include data warehousing and mining and distributed database systems.
  • Authors: Dr. Pragnayaban Mishra is a professor and head of information technology with research areas spanning data warehousing, mining, distributed systems, databases, operating systems, and neural networks.The passage also states that he guides five PhD scholars.
  • Authors: Mrs. Rasmita Panigrahi is a lecturer pursuing an MTech in computer science, with interests in data warehousing, mining, distributed databases, design, algorithms, and cryptography.The passage identifies her affiliation with Gandhi Institute of Engineering and Technology.
Loading 1211.5723v1…