Source-linked AI summary

Service Level Agreement (SLA) in Utility Computing Systems

Linlin Wu, Rajkumar Buyya

arXiv:1010.2881v1cs.DC

TL;DR

Utility-computing research lacks an overall classification of extensive SLA languages, frameworks, and management work, despite SLAs’ role in meeting pre-agreed service expectations. This chapter surveys SLA creation, management, lifecycle models, and Grid and Cloud use cases, identifying ongoing issues and future research challenges. It concludes that scalability, dynamic changes, heterogeneity, automation, multiple QoS parameters, and cross-domain support remain open challenges.

  • Problem

    Utility-computing SLA research lacks an overall classification, while management must address autonomy and trade-offs among multiple QoS parameters.

  • Method

    The chapter surveys SLA implementation and management challenges, solutions, frameworks, languages, lifecycle models, and Grid and Cloud use cases.

  • Results

    The survey identifies lifecycle models, existing SLA frameworks and languages, use-case realization patterns, and key design issues in utility-computing systems.

  • Takeaways & Limitations

    The analysis identifies four goals: customer-driven service management, computational risk management, market-based resource management, and autonomic resource management.

  • Takeaways & Limitations

    SLA management remains a rapidly moving target with open challenges in scalability, dynamic changes, heterogeneity, automation, multiple QoS parameters, and cross-domain support.

Abstract

from arXiv · show

In recent years, extensive research has been conducted in the area of Service Level Agreement (SLA) for utility computing systems. An SLA is a formal contract used to guarantee that consumers' service quality expectation can be achieved. In utility computing systems, the level of customer satisfaction is crucial, making SLAs significantly important in these environments. Fundamental issue is the management of SLAs, including SLA autonomy management or trade off among multiple Quality of Service (QoS) parameters. Many SLA languages and frameworks have been developed as solutions; however, there is no overall classification for these extensive works. Therefore, the aim of this chapter is to present a comprehensive survey of how SLAs are created, managed and used in utility computing environment. We discuss existing use cases from Grid and Cloud computing systems to identify the level of SLA realization in state-of-art systems and emerging challenges for future research.

1. INTRODUCTION

Utility computing delivers on-demand, pay-per-use services, making SLAs important for defining expectations, managing service quality, and guiding provider–consumer relationships. The chapter surveys SLA creation, management, lifecycle, use cases, and unresolved design issues in Grid and Cloud systems.

  • Utility computing provides subscription-oriented services on demand, allowing users to outsource jobs and pay for usage instead of maintaining infrastructure.
  • SLAs specify parties, pricing, resource properties, service expectations, and obligations within utility-computing transactions.
  • Clearly defined SLAs can improve customer satisfaction, service quality, and relationships by aligning requirements, KPIs, monitoring, and penalties.
  • Effective SLA realization requires a lifecycle spanning creation, operation, and removal, or six detailed steps from provider discovery through penalty enforcement.
  • Grid and Cloud systems use SLAs amid changing resource conditions, while monitoring, assurance, autonomy, and multi-QoS management remain significant challenges.
  • The chapter surveys existing designs and use cases to identify key factors and issues for extending SLA frameworks and implementing enhanced SLA-oriented management systems.

2. UTILITY ARCHITECTURE AND SLA FOUNDATIONS

The chapter presents utility-system architecture, SLA definitions and components, and two lifecycle models. It treats the six-step lifecycle as the more detailed characterization because it includes negotiation and violation control.

  • 2.1. Utility Architecture: Utility architecture places users or brokers above admission control, SLA management, and resource or service providers.SLA management handles resource allocation, while the Service Request Examiner decides whether requests are accepted or rejected.
  • 2.1. Utility Architecture: SLA management includes discovery, negotiation, pricing, scheduling, monitoring, enforcement, dispatching, and accounting components.Pricing helps manage supply and demand and prioritize allocations; monitoring tracks resources and request execution, while enforcement handles contract violations.
  • 2.2. SLA Definitions: SLA definitions in information technology vary by area, but commonly describe provider capability, consumer performance targets, availability guarantees, and measurement mechanisms.
  • 2.3. SLA Components: SLA components include purpose, restrictions, validity period, scope, parties, service-level objectives, indicators, penalties, optional services, and administration.Indicators include availability, performance, and reliability; penalties apply when SLOs or performance measurements are not achieved.
  • 2.4. SLA Lifecycle: The high-level lifecycle has creation, operation, and removal phases, covering provider selection, SLA access during operation, and termination with configuration removal.
  • 2.4. SLA Lifecycle: The detailed lifecycle comprises discovery, SLA definition, agreement establishment, violation monitoring, termination, and penalty enforcement.Its definition stage can include negotiation, while enforcement invokes penalty clauses when contract terms are violated.
  • 2.4. SLA Lifecycle: The six-step model is presented as a better characterization because it explicitly includes negotiation or renegotiation and violation control.Negotiation exchanges contract messages to reach mutual agreement, producing a new SLA.

3. SLA IN UTILITY COMPUTING SYSTEMS

Utility-computing SLAs must account for customer priorities and changing operational conditions. The chapter therefore emphasizes dynamic QoS requirements, including reliability and trust/security, in SLA management.

  • Service providers use SLAs to define user-required service parameters and obtain feedback about how users value service requests.
  • Commercial services require QoS parameters such as reliability and trust/security in addition to standard service requirements.
  • QoS requirements need dynamic updates because business operations and operating environments continuously change.

3.1. SLA Management in Utility Computing Systems

SLA management in utility computing spans resource discovery, agreement formation, monitoring, termination, and penalty enforcement. Dynamic resources, heterogeneous policies, negotiation differences, and fairness make these lifecycle stages challenging.

  • Utility environments require efficient resource discovery across geographically distributed resources with heterogeneous administrative policies and dynamic membership.
  • SLA service terms cover QoS parameters, provider delivery ability, workload performance targets, availability and performance bounds, reporting, cost, renegotiation data, and penalties.
  • Negotiation is complicated by differing protocols, service definitions, and the need for unambiguous, parameterized, context-specific service descriptions.
  • Automated negotiation and self-renegotiation after failures remain insufficiently developed, while SLA templates have difficulty reflecting component quality.
  • SLA violation monitoring must determine responsibility, ensure fairness, and define violation boundaries for deciding whether SLOs are achieved.
  • Violation provisioning may be All-or-Nothing, Partial, or Weighted Partial, depending on whether all SLOs, mandatory SLOs, or weighted thresholds must be satisfied.
  • Penalty enforcement requires comprehensive and fair clauses; linear models used in simple contexts exhibit poor performance, leaving model selection open.

3.2. Solutions for SLA Management in Utility Computing Systems

The chapter surveys six SLA languages and frameworks and related Grid systems as solutions for specifying, negotiating, monitoring, and enforcing agreements. These systems address different lifecycle needs but retain interoperability, semantic, scalability, and penalty-design limitations.

  • Six SLA management languages and frameworks are analyzed because they support multiple steps of the SLA lifecycle.
  • SLA languages support preparation, automated negotiation, service adaptation, and reasoning about service composition; WS-Agreement and WSLA are the most widely used.
  • WS-Agreement uses XML templates and request-response interactions for agreement establishment, provider discovery, status exposure, and dynamic violation management.
  • WSLA provides an XML-based language and runtime architecture for measuring QoS, monitoring violations, reporting them, and separating monitoring clauses from contractual terms.
  • WSOL, SLAng, QML, and QuO respectively emphasize reusable service offers, industry-specific vocabularies, metric type systems, and proxy-based QoS adaptation.
  • MDS uses a static information-server relationship, while VIRD improves scalable hierarchical discovery but does not address heterogeneity or autonomous administration.
  • Meta-negotiation documents address incompatible protocols, languages, and prerequisites that otherwise prevent consumers and providers from negotiating across different standards.
  • Existing resource-management work includes risk-to-penalty assessment and negotiation mechanisms, but one auction solution supports price alone rather than multiple QoS dimensions.

4. SLA USE CASES IN UTILITY COMPUTING SYSTEMS

Grid and Cloud use cases show SLA support for collaboration, risk assessment, renegotiation, layered service delivery, and provider-defined commercial agreements. Cloud systems commonly simplify lifecycle activation through predefined SLAs and external monitoring tools.

  • Utility computing offers on-demand IT capabilities under usage-based pricing, with Grid and Cloud as two approaches for exploiting idle data-center capacity.
  • Grid SLA projects are classified into business collaboration, risk assessment, and renegotiation supporting dynamic changes.
  • GRIA supports cross-organizational collaboration by discovering resources, assigning them through SLAs, charging for usage, and monitoring services.
  • AssessGrid maps failure risk to penalty fees, helping providers select SLA offers and users judge acceptable costs and penalties.
  • A proposed WS-Agreement extension enables runtime renegotiation, but it does not adapt agreements to dynamic operational and environmental changes after establishment.
  • Cloud computing provides on-demand infrastructure, platform, and application services, with SLAs negotiated between providers and consumers.
  • Cloud architecture layers physical resources, core middleware, user-level middleware, and applications, with dynamic SLA management among core middleware services.
  • Published Amazon and Microsoft SLA documents were used to summarize industry parameters and characterize systems against a six-step SLA lifecycle.

5. ONGOING WORKS

Ongoing SLA research addresses scalable, adaptive management across heterogeneous utility-computing environments, while several automation, negotiation, and cross-domain challenges remain open.

  • SLA management must reliably provision services, monitor violations, and detect performance degradation during execution.
  • Scalable, automatic frameworks must adapt to dynamic environmental changes while considering multiple QoS parameters and heterogeneous resources.
  • Multiple-dimensional auctions are important for utility computing, because one-dimensional auction mechanisms cannot handle them.
  • Future SLA management should retain consumer involvement and dynamically address price, deadline, reliability, and trust/security requirements.
  • Different negotiation protocols constrain SLA establishment, modification, and negotiation across distinct administrative domains.

6. SUMMARY

The chapter surveys SLA management and use in utility computing, organizing its findings around lifecycle phases, management goals, implementation challenges, and unresolved research problems.

  • The chapter surveys SLA management issues, solutions, and uses in utility computing systems.
  • A comprehensive six-step lifecycle extends the three-phase model by characterizing SLA violations for pay-as-you-go services.
  • The analysis identifies four goals: customer-driven management, computational risk management, market-based resource management, and autonomic adaptation.
  • Lifecycle-based analysis discusses scalability, dynamic change, heterogeneity, negotiation, SLA languages, monitoring responsibility, and third-party monitoring.
  • Automatic negotiation, problem resolution, and cause analysis remain open challenges requiring further investigation.
  • Future research must address scalability, dynamic environments, heterogeneity, automation, multiple QoS parameters, and cross-domain suitability.

ADDITIONAL READING SECTION

The additional reading section lists related works on Grid and utility computing, covering SLA negotiation, reservations, superscheduling, resource allocation, middleware, and market-oriented systems.

  • Related work includes a chapter on Service Level Agreements in the Grid Environment.
  • A related chapter examines SLAs, negotiation, and potential problems.
  • Additional studies address SLA-based advance reservations with flexible and adaptive time QoS parameters and coordinated superscheduling in computational Grids.
  • Other references cover market-oriented Grids, the Gridbus Toolkit, SLA-based resource management, and Gridbus middleware.

AUTHORS PROFILE

The author profiles identify Linlin Wu as a CLOUDS Laboratory PhD candidate and Rajkumar Buyya as a University of Melbourne professor and CLOUDS Laboratory director.

  • Linlin Wu is a PhD candidate supervised by Rajkumar Buyya at the University of Melbourne’s CLOUDS Laboratory.
  • Wu’s research interests include SLA, QoS measurement, resource allocation, and market-oriented Cloud computing.
  • Rajkumar Buyya is Professor of Computer Science and Software Engineering and director of the CLOUDS Laboratory at the University of Melbourne.
Loading 1010.2881v1…