Source-linked AI summary
Big Data Computing and Clouds: Trends and Future Directions
Marcos D. Assuncao, Rodrigo N. Calheiros, Silvia Bianchi, Marco A. S. Netto, Rajkumar Buyya
TL;DR
Organisations need scalable ways to manage and analyse growing, heterogeneous data, while Cloud analytics still has technical and organisational gaps. This paper surveys Cloud-supported Big Data analytics across data management, modelling, visualisation, interaction, and business models. It concludes that tool selection must match analytics requirements and that future work should improve interoperability and Cloud elasticity while retaining human involvement where needed.
Problem
Growing and heterogeneous Big Data creates challenges for data management, processing, modelling, interaction, and organisational adoption of Cloud analytics.
Method
The paper surveys Cloud-supported Big Data analytics solutions across data management, model building and scoring, visualisation, user interaction, and business models.
Results
The survey finds many Cloud Big Data solutions, but their variety can overwhelm inexperienced users and requires matching tools to analytics requirements.
Takeaways & Limitations
Future work should develop standards and APIs for switching among solutions and improve use of Cloud elasticity for analytics.
Takeaways & Limitations
High-level Cloud analytics services may still require human analysts because machine learning and Big Data analysis cannot easily replace human expertise in some scenarios.
Abstract
from arXiv · showhide
This paper discusses approaches and environments for carrying out analytics on Clouds for Big Data applications. It revolves around four important areas of analytics and Big Data, namely (i) data management and supporting architectures; (ii) model development and scoring; (iii) visualisation and user interaction; and (iv) business models. Through a detailed survey, we identify possible gaps in technology and provide recommendations for the research community on future directions on Cloud-supported Big Data computing and analytics solutions.
1. Introduction
Big Data analytics can create organisational value from private and public data, but implementing it remains complex and resource-intensive. The paper surveys Cloud-supported analytics approaches and identifies technology gaps and future research directions.
- Organisations generate vast amounts of data, making data management and insight extraction challenging but important for competitive advantage.
- Putting analytics and Big Data into practice requires expensive software, substantial computing infrastructure, consulting, and organisational effort.
- Clouds offer pay-as-you-go resources, elasticity, availability, and potential cost reduction for analytics applications.
- Cloud analytics still faces challenges that must be addressed before Clouds become ideal platforms for scalable analytics.
- The paper surveys Cloud analytics approaches, environments, and technologies, covering technical and non-technical issues and proposing research gaps and recommendations.
2. Background and Methodology
The paper frames Big Data analytics as a workflow that prepares heterogeneous data, builds and validates models, scores incoming data, and supports descriptive, predictive, and prescriptive analysis. It surveys Cloud solutions across these stages while highlighting technical requirements and business challenges.
- Big Data combines data from sources such as databases, streams, marts, warehouses, sensors, and social networks, creating challenges for existing infrastructure.
- Analytics workflows integrate, clean, and filter source data before training and parameter estimation, then validate the resulting model.
- Model scoring applies validated models to arriving data to generate predictions, prescriptions, and recommendations.
- Descriptive analytics models past behaviour, predictive analytics estimates future outcomes, and prescriptive analytics assesses actions against objectives and constraints.
- Cloud-hosted analytics can be consumed pay-as-you-go, but requires attention to data management, model tuning, privacy, data quality, and data currency.
- The survey examines Cloud data management, model development, visualisation, interaction, service structures, service-level agreements, and business models.
3. Data Management
Cloud Big Data management must address diverse deployment models, specialised resources, varied data characteristics, storage locality, integration, and interactive or real-time processing demands.
- Deployment models: Cloud analytics spans private, public, hybrid, and mixed data-model deployment scenarios, each balancing control, cost, elasticity, security, or provider management.Private Clouds maximise control; public Clouds offer shared low-cost resources; hybrid Clouds add public capacity to private environments.
- Deployment models: Analytics Cloud services may need to manage specialised data and domain experts alongside conventional computing resources to achieve elasticity and economies of scale.The paper distinguishes these requirements from ordinary Cloud services and calls for proper allocation and utilisation of specialised resources.
- Data variety and velocity: Big Data management must accommodate variety, velocity, volume, veracity, and value, while recognising that veracity and related data-quality and provenance issues warrant focused study.Variety includes heterogeneous public and application-specific sources, while velocity ranges from batch processing to continuous real-time analysis.
- Data storage: Large-scale analytics makes data locality important because transferring datasets to computing units can be prohibitive; MapReduce and HDFS exploit locality through partitioning and replication.HDFS also reduces failure impact by replicating datasets across configurable numbers of nodes.
- Data integration solutions: Cloud data warehouses face data-integration and new-source challenges, motivating standard formats, interfaces, composable spaces, and workflow-oriented analytics services.These approaches also support gradual migration from on-premise analytics to Cloud-provided infrastructure.
- Data processing and resource management: Processing systems must support stringent near-time or stream deadlines and emerging small, short, highly interactive jobs, exposing limits in conventional MapReduce-based analytics.Continuous analytics services combine stream processing with relational data access through SQL-like interfaces and cycle-based execution.
4. Model Building and Scoring
Cloud analytics uses storage and data services to build and score models, while research addresses scalable model development, interoperability, and abstractions for trained systems.
- Model Building and Scoring: Cloud analytics must build models for forecasts and prescriptions and test them against new data through model building and scoring.These activities can be offloaded to Cloud providers, with some machine-learning algorithms parallelised.
- Cloud Model Deployment: Predictive models can be deployed on Amazon EC2 and exposed through Web Services interfaces using PMML.PMML is also advocated as an exchange language for predictive models.
- Cloud Model Deployment: Zementis supports model analysis and building on customer premises or as SaaS using IaaS platforms such as Amazon EC2 and IBM SmartCloud Enterprise.
- Cloud Model Deployment: Google Prediction API lets users train models from submitted data, share models, and predict numeric values or categories for new items.
- Abstractions: The Hazy project targets programming and infrastructure abstractions to simplify assembling existing solutions into trained systems.The stated goal is to ease the complexity of building systems such as Watson, Siri, and Google Knowledge Graph.
- Research Challenges: Rapid Cloud elasticity and increasing data volumes create a need for timely model building and scoring, alongside standards that support competing analytics services without vendor lock-in.Standards APIs and formats would let customers choose providers based on cost and performance.
5. Visualisation and User Interaction
Big Data visualisation and interaction require scalable processing, suitable representations, and interfaces that reduce the delays of batch-oriented Cloud analytics workflows.
- Visualisation Requirements: Visualisation tools should account for data quality, presentation, and data volume while supporting descriptive, predictive, and prescriptive analytics.Selecting visualisation types according to displayed data can improve display and performance.
- User Interaction: Cloud analytics commonly follows a batch-job model in which users submit jobs, wait for completion, and inspect downloaded samples afterward.This back-and-forth is poorly supported by the Cloud and motivates better interactive interfaces.
- User Interaction: Incremental query-result representations were robust enough for interviewed analysts to abandon, refine, or reformulate queries.
- User Interaction: Interfaces have connected familiar tools such as spreadsheets to Cloud analytics backends, including Excel integration with Daytona and Azure infrastructure.
- Visualisation Systems: Visualisation systems offer selectable methods, data navigation, peer discussion, dashboards, widgets, charts, and multi-source aggregation.Examples include ManyEyes and domain-specific tools for climate modelling and computing-network management.
- Research Challenges: Real-time Big Data visualisation is constrained by analytics complexity and display capacity, motivating approximate processing and large-scale visualisation environments.Proposed techniques include reduced precision, coarse processing, reduced convergence, and data-scale confinement.
6. Business Models and Non-Technical Challenges
Cloud analytics business models range from shared platforms and full-stack services to hosted models, but must address replication, pricing, isolation, and the balance between generality and usefulness.
- Business Models: Analytics providers have been studied as a way to move customised analytics solutions from customer premises to Cloud delivery.
- Business Models: Cloud analytics services can be organised as shared platforms, end-to-end full stacks, or hosted model-scoring services.These models target organisations with multiple analytics departments, limited analytical expertise, or insufficient data for good predictions, respectively.
- Business Models: Analytics as a Service provides on-demand analytics, while Model as a Service offers models as building blocks for analytics solutions.
- Non-Technical Challenges: Multi-tenancy proposals seek lower costs but introduce challenges in isolating analytical artefacts, pricing services, and defining Service Level Agreements.
- Service Examples: Managed analytics services such as IVOCA aim to reduce time to insight and improve repeatability in CRM analytics.
- Non-Technical Challenges: Cloud analytics must balance generality and usefulness because customer-specific solutions and models often require updates for new data, complicating replicability.Prior work also discusses difficulty replicating text-analytics activities and proposes pathways linking business objectives to analytical flows.
- Service Ecosystems: An analytics ecosystem can layer analytics services over Data as a Service, integrating public and private data through providers, integrators, aggregators, and clients.
7. Other Challenges
Other Cloud analytics challenges include human and organisational constraints, cost estimation, weak ingestion, limited interactivity, emerging mobile and sensor requirements, and poorly defined service contracts.
- Human and Organisational Challenges: Human expertise cannot easily be replaced by Cloud-delivered machine learning and Big Data analysis, so analysts may need to remain in the loop.Management must also help analysts gain insights and support quicker decisions.
- Cost and Deployment: Cloud application profiling must estimate data-transfer, virtual-machine, and execution costs before targeting a platform, but current offerings make this estimation non-trivial.
- Development and Validation: Data ingestion is often weak, while debugging and validating Cloud analytics solutions are challenging and tedious.
- Development and Validation: Current Cloud environments lack interactive analysis workflows that iteratively refine query answers and give users more control of processing.Such techniques are intended to reduce time to insight and include analysts in the loop.
- Human and Organisational Challenges: Corporate analytics is hindered by inadequate staffing and skills, insufficient business support, and software problems, which Cloud delivery can exacerbate.
- Emerging Requirements: Advanced visualisation and analytics for unstructured, large, and streaming data are expected to become important to organisations.
- Emerging Requirements: BI&A 3.0 requires mobile, location-aware, and context-aware techniques for large-scale mobile and sensor data, many of which remain undeveloped.It also requires integrating multiple data sources for Cloud processing.
- Service Contracts: AaaS and BDaaS lack well-defined contracts for result quality, input-data reliability, execution times, methods, and responsible experts.
8. Summary and Conclusions
Cloud-supported Big Data computing offers many analytics solutions, but their diversity and complexity make requirement-driven tool selection and expert involvement essential. The survey identifies Cloud analytics as an emerging area, highlights service-based and elastic approaches, and points to standards and APIs as recurring future priorities.
- Cloud resources can be scaled up and down with demand, helping address the resource requirements of Big Data analytics.Cloud computing provides resources on demand with costs proportional to actual usage.
- The survey classifies Cloud-supported analytics research across data management, model building and scoring, and visualisation and user interactions, followed by business-model and non-technical analysis.For each technical area, it reviews ongoing work and discusses open challenges.
- Big Data Cloud computing offers many solutions because analytics requirements and data characteristics vary, but this diversity can overwhelm inexperienced users.Requirements should guide the selection of appropriate tools across descriptive, predictive, and prescriptive analytics and different data properties.
- Analytics remains complex and requires expertise in data preparation, method selection, and result analysis; Analytics as a Service or Big Data as a Service may offer an alternative to in-house work.The suitability of these providers depends on the complexity and costs of the tasks involved.
- Cloud computing supports Big Data through infrastructure, tools, and service-based business models, including AaaS and BDaaS, but these models involve greater customer and provider participation.This involvement creates challenges beyond those of traditional infrastructure, platform, and software services.
- Recurring future directions include standards and APIs for switching among solutions and better use of Cloud elasticity.The survey also identifies expressive languages as part of exploiting elastic Cloud infrastructure.