Showing posts with label Storage. Show all posts
Showing posts with label Storage. Show all posts

Monday, February 4, 2013

Three reasons why hybrid storage shines for mainstream workloads

Rob Commins, vice president of marketing, Tegile Systems, says:

If you’ve read stories about the “death of the hard drive,” you may have been misled into believing that there is no longer a place in the enterprise storage market for what many storage architects fondly refer to as “spinning rust.” It is certainly the case that storage is undergoing a major transformation, but it’s clear that smart enterprises are not completely abandoning what is tried and true technology just yet.  Additionally, there are vendors actively working with these enterprises and providing them with solutions which intelligently leverage traditional, high capacity hard drives as a cornerstone component of their next-generation storage systems. When combined with more modern, speedy solid state storage devices, new life is breathed into traditional hard drives and customers that purchase these hybrid HDD/SSD storage arrays are finding that the outcomes are exactly what the business needs to continue to forge ahead.

As is the case for everything, there are certainly fringe use cases that call for something other than hybrid storage. Pure archiving, for example, generally calls for large capacity hard disk drives while very high end analytics might call for an all-flash solution that can deliver a million IOPS or more in a single chassis.

For mainstream workloads — in other words, the other 95% of the market — hybrid storage provides the perfect solution for mixed I/O workloads. Here are three reasons that hybrid shines for these mainstream environments.

Balanced cost vs. performance

CIOs today demand the best possible return on their investment and they’re not willing to simply throw money at narrowly purposed and over-engineered sub-systems to get the kind of performance they need to run their workloads. A more flexible solution is required.

Enter hybrid.

Hybrid storage arrays provide CIOs with the best of both worlds when it comes to balancing real world capacity needs with real world performance needs. Even though many all-flash storage vendors make the news by demonstrating a million+ IOPS from their arrays, most CIOs simply don’t need that kind of performance and aren’t willing to pay the very high cost per gigabyte (capacity) that accompanies the low cost per IOPS (performance) from these arrays.
By buying hybrid, CIOs are seeing $/GB prices that approach that of hard disk drive-based storage while paying $/IOPS prices that simply isn’t possible to achieve with HDDs or SSDs alone. Only when HDDs and SSDs are brought together into one is it possible to truly balance capacity and performance needs.

Simplified technology environment

Another trend that is seeing momentum revolves around simplifying the technology environment. Storage systems haven’t generally been considered the easiest IT infrastructure components to manage and over the years, the whole paradigm has become even more complex as IT departments deploy new kinds of workloads with wildly variant I/O patterns.

In considering the history of VDI, for example, CIOs that were early adopters of this technology often found that they needed to deploy completely separate storage environments in order to contend with the unique challenges wrought by the desire to go thin. In other cases, CIOs have supported environments that required complex masses of tiered storage with storage administrators carefully carving out new LUNs based on what they hoped were correct workload specifications.

With a hybrid solution, IT departments can jettison the tiers and have just one single tier of storage: Fast. And again, the pure performance of a hybrid solution doesn’t come with a capacity trade off. Organizations get it all in one box.

No compromise

It boils down to this: With a hybrid storage solution for mainstream workloads, CIOs don’t have to compromise. They get capacity and performance, but with almost all of the solutions that have hit the market in recent years, they also get enterprise grade features, such as de-duplication (which further improves the capacity of the solution and, for some vendors, such as Tegile, actually improves the overall performance of the workloads hosted on the array), compression, scalability (up and/or out depending on the vendor) and, with some vendors, such as, again, Tegile, a broad suite of communications protocols that allow the hybrid array to simply drop into just about any modern technology environment with little to no hassle required.

For these reasons and many more, hybrid storage arrays will be the sweet spot in the storage market for the foreseeable future for just about any mainstream workload need.

About Rob Commins
Rob Commins is vice president of marketing, Tegile Systems. He has been instrumental in the success of some of the storage industry’s most interesting companies over the past twenty years. As Vice President of Marketing at Tegile, he leads the company’s marketing strategy, go to market and demand generation activities, as well as competitive analysis. Rob comes to Tegile from HP/3PAR, where he led the product marketing team through several product launches and 3X customer growth over three quarters. Rob also managed much of the functional marketing and operations integration after Hewlett Packard acquired 3PAR. At Pillar Data Systems, he was at the forefront of converged NAS/SAN storage systems and application-aware QoS in mid-range storage. Rob is also a veteran of StorageWay, one of the first storage services providers that launched cloud services.

Tuesday, December 4, 2012

Outlook and QuickBooks: What You Should Know BEFORE Your Computer Crashes


- Jamie Brenzel, CEO at KineticD, says:

It's a worst case scenario - the network is underwater and the entire organization’s critical applications and associated data have disappeared. It is every business owner’s fear. Whether it’s due to theft, a hard drive crash or a hurricane, the loss of company data can cripple progress, leaving businesses stranded and holding the proverbial bag.

Business owners often don’t think about the procedures necessary to backup data, much less the recovery process involved in the event of a system failure. Unfortunately, what many don’t realize is that the stalwart back up methods of the past such as tape, DVD’s or CD’s, have become painfully old fashioned in today’s data-centric world.

While many small-to-mid-sized businesses (SMBs) think ALL of their data is being captured and stored in a safe place, the reality is that very often they are wrong. The truth is that many mission critical programs are not being backed up because, when left open, they are outside the parameters of standard backup procedures.

The golden question all business owners should ask is simple: “If my computer disappeared tomorrow, could my business survive without past emails, contacts, work schedules and accounting data?” If the answer is no, then it’s time to take a closer look at your backup and data recovery procedures.

The Dilemma
Critical non-standard databases, such as QuickBooks and Outlook along with other accounting software and legacy applications, as well as databases such as MS Access and MySQL, may not be VSS-aware. This means they lack the standard that allows files or databases to be backed up when in use. This lack of functionality leaves many SMBs and IT managers empty-handed in the event of a system failure or disaster.

QuickBooks and Outlook data is infamous for being left out of standard backup routine practices. The ongoing dilemma of backing up this critical data continues to be the fact that most employees, even CEO’s, forget to shut down these programs when leaving for the day. Unless special precautions are taken, it is likely that when attempting to restore these critical data files they will be outdated… or worse, corrupted. This is true of many mission-critical programs that keep businesses running.

Outlook is a critical application for businesses of all sizes. If it is not closed, many backup tools are unable to access the ever-important .PST files, leaving companies open to a potential disaster should the program crash. The .PST contains all of your Outlook data, including emails, contacts, calendar events, notes and schedules; and file backup doesn’t necessarily happen automatically. Regular backup of the .PST file is critical for restoring your most recent data.

For those of us who are email hoarders, keeping all email correspondence in Outlook can cause the .PST to reach gigabyte proportions very quickly. These large and unruly files can produce unexpected problems when backing them up.  Many backup systems can “time out” or even worse, demand an extra tape or CD to complete the backup routine when your office is closed and employees are gone for the day.

A good starting point is to look for a Microsoft Certified backup vendor that follows the best practices set by Microsoft. Selecting backup software that doesn’t conflict with Windows or other low-level drivers, such as antivirus programs and software firewalls, should be a key element in identifying a backup vendor.

Our personal data and financial records rank high on the list of importance. This is especially true for business owners. QuickBooks, the popular accounting software is another application many SMBs can’t do without. The absence of an Application Programming Interface (API) presents ongoing challenges for protecting corporate financial data. An API allows third party companies to integrate special functions, such as requests for a data dump for easy backup. Currently, there is no safe method to programmatically dump your QuickBooks data to a secure place for easy backup at the end of the day, and without an API, developers can’t create one for QuickBooks.

Scheduling QuickBooks to back up automatically to a specific folder is recommended, which allows your data to be verified regularly. From there, schedule your third party backup software to back up to that folder. This ensures that QuickBooks is in a good state before your backup software runs. Not following these best practices could lead to a corrupted QuickBooks database file.

QuickBooks files that your company must make sure are backed up:
.QBW           =    quickbooks Primary Data File
.QBB            =    quickbooks Backup File
.QBA           =    quickbooks Accountant’s Copy File
(May also be a .AIF [Accountant’s Import File])
.QBA.TLG   =    Transaction log file (for the accountant’s review copy)
.QBM           =    quickbooks Portable Company File (for version 2006 and above).
.QBI             =    quickbooks Crash Roll Back File
.QBX           =    quickbooks Accountant Transfer File
.QDT            =    quickbooks United Kingdom Accountant Data File

The backup challenges of these programs are unique; so considering the alternatives that address this lack of functionality is important. Some data backup companies provide solutions that include the ability to continuously back up your programs and data, no matter where that device is located and even while the programs are open and in use.

Data Security Drill
If you are questioning how secure your data is, you should test it. Just like any safety drill, the best way to make sure is to put your backup method to the test. If possible, locate a spare computer, wipe the drive and attempt to do a complete restore of company applications and associated data. Then ask the following questions:
  • Are you able to restore your backup files in one simple step?
  • Do you have easy access to those applications and have you kept track of the associated upgrades?
  • Do you need to re-install all of your applications individually?
  • Does your Outlook, QuickBooks and other proprietary data install easily?
  • Do you end up with the most current data when you have completed the restoration?
Keep track of the time it takes you to complete the restoration from beginning to end, then multiply that by the number of computers you have within your organization. Many SMBs are unaware that the secret of data recovery is the time it takes to restore the data, which is generally governed by your company’s Internet bandwidth restrictions. When your computers are down, every minute counts.

Conclusion
Data backup is as important to your business as an insurance policy. Computers can be replaced if lost in a disaster, or stolen during a break-in. But if you don’t have a reliable, secure method of restoring your data, all the money in the world isn’t going to bring back those missing files.

Ideally there would be no need to back up your company’s data; computers and hard drives would last forever, outside threats like malicious hackers and natural disasters would be non-existent and employees would never forget to save a file, or close a program. Unfortunately this perfect world does not exist. For this reason, businesses of all types and sizes should take the necessary steps to protect their data at all times if they wish to remain in business after the fact.

About the Author
Jamie brings over 15 years of experience in investment banking and entrepreneurial startups to his role as CEO of KineticD. Jamie holds an Honours Bachelor of Arts degree in Politics and Philosophy from the University of Western Ontario.

Thursday, November 29, 2012

Don’t Spring an Information Leak


- Trevor Daughney, director of Product Marketing with Symantec Corp., says:

You may have seen a cartoon in which a character is trying to plug a hole in a dam with his finger. As he plugs one, another leak is sprung, and he sticks another finger in that one. Another leak is plugged by his toe, until he can no longer hold back the water and the dam bursts. This is comparable to the situation in many of today’s data centers, which are becoming so overrun with information that it can overwhelm an organization’s resources.

The 2012 Symantec State of Information Survey was created to examine how businesses are dealing with the massive increase in information today, especially as cloud computing and mobile devices change the way information is stored and accessed. The survey uncovered key challenges facing organizations today. One challenge is the large amount of duplicate information being created – 42 percent of all business information, in fact. This leads to inefficient storage management and other challenges, such as knowing where to find the most current version of a file. In a related finding, respondents reported that more than 40 percent of their information is hard to find.

Another challenge arises because of the difficulty in giving all information equal weight. 31 percent of organizations reported not knowing how important their stored information is. With the constant increase in threats to information today, that can lead to increased risk of data loss or theft because of insufficient protection for sensitive information. It may also lead to retention of files longer than necessary, which can lead to problems in case of litigation or regulatory issues. In fact, nearly one-third of organizations surveyed reported that they had missed compliance requirements within the last year.

Despite these challenges, it is possible for businesses to combat the threat of drowning in their information. Creating an information governance plan can greatly improve an organization’s security and efficiency, but it must be supported by the entire company. The plan should include a combination of policies and processes to complement the technologies you are using to protect your information. Even the best protection will fail if it is not applied consistently. Symantec recommends following three steps to take control of your information.
  • First, receive top-down support for the initiative. C-level support is essential for any information governance program to succeed. Managerial oversight is necessary in order to coordinate information protection with the overall goals of the business. It will also help break down communication barriers between business entities.
  • Ensure that the program consists of specific projects with concrete goals. Different aspects of governance must be addressed, including security, compliance and eDiscovery. Intelligent management solutions such as archiving, data loss prevention and deduplication provide cost-effective protection to complement policies.
  • One of the most important steps to achieving data safety is to know what you are protecting. Commit the necessary resources to categorizing the information you are currently storing, to effectively focus protection efforts.

By understanding how your information relates to your organization, and engaging employees and solutions in a specific plan to keep it safe, you can keep your organization from being flooded with problems, retaining a controlled reservoir of valuable information that will help keep your business moving forward.

Friday, November 23, 2012

Beyond RAID: Wide Area Storage for Large-Scale Disk Archives


Janet Lafleur, Product Marketing Manager at Quantum, says:

RAID was built to protect data from disk failures, but was it built to scale?  When you consider what happens when disk capacities grow to the multi-terabyte level and how many disks are needed to meet the growing demand for big data and media files, the answer is clearly “No.” 

When a 3TB or larger drive fails, it can require over 24 hours to rebuild the data on a replacement drive.  Until the failed drive is detected, replaced and the rebuild is completed, the storage array is left vulnerable to data loss from another drive failure, experiences degraded performance or both.  Add to that the regular supervision needed to replicate data across RAID arrays to deliver large files to remote users and for disaster recovery, and the operational challenges soar along with the data volumes.

For those who manage large repositories of large files shared by a large number of users, scale-out NAS storage built on RAID arrays has been the go-to technology for years. But when the storage demand scales toward and beyond the petabyte level, a new architecture is needed for disk-based archives, one with extreme scalability and durability that reduces both capital and operating expenses.  That new architecture is wide area storage.

Wide area storage combines next generation dispersed object storage with file system technologies in a new approach to archiving that overcomes these limitations and inefficiencies. Data repositories built on wide area storage are extremely scalable, durable and easy to maintain--allowing data to be stored forever on disk without business-halting service interruptions or painful data migrations.

Unlike RAID, wide area storage uses fountain coding, a type of forward error correction algorithm developed for communication over unreliable networks. The fountain code algorithm encodes the data into a set of equations in such a way that fewer than the full set is needed to reconstruct the original data. The equations are then dispersed across the storage which can be in multiple geographic sites.

The higher redundancy gives wide area storage 15 nines of durability—far greater than RAID—which means there’s no rush to replace failed drives and initiate rebuilds. Maintenance can be performed on a scheduled, not rushed basis.

Unlike RAID, wide area storage can mix and match drives from different technologies within the storage system. This makes the upgrade process as simple as removing a set of drives from active service and replacing with new ones. The wide area storage will re-calculate and disperse the data on the new drives and continue to run throughout the upgrade process.

The result: disk storage that scales indefinitely, preserves data indefinitely, and reduces maintenance for long-term, large-scale archives.

Friday, November 16, 2012

IT Efficiency with Deduplication Technology







Wayne Salpietro, director of marketing at Permabit, says:

In today’s information technology world, it is time to adopt deduplication across the entire IT spectrum from applications to operating environments across all of the storage media (flash, SSD, HDD) and implementations (primary, 2nd tiers, archiving and the cloud) and, last but not least, in all virtual environments. You might be asking “why now?” There are several trends that are intersecting today that make it a necessity. Let’s review them:

Data growth is rampant, creating a data deluge that is nearly impossible to afford and manage.

Information growth is yielding larger and larger data stores which have spawned a new phrase “Big Data” and big data is being used by analytics engines to identify trends that may lead to differentiation and incremental revenue.

Storage transformation is occurring as solid state devices rapidly gain traction because their performance makes them a must have addition to existing spinning disk storage. As solid state technology (disk, flash and cache) closes the “cost gap” with spinning disk, the performance gains will be so compelling that rampant adoption will occur and storage will be changed forever!

Virtualization is delivering more optimization of user functionality and ease-of-use that it too is being adopted across enterprises to deliver information transparently enabling faster –better – more efficient customer interactions and increased revenue generation.

Consumerization of technology is being added to an increasing virtual desktop infrastructure (VDI) adoption as “once personal devices” are being integrated into the enterprise to harness worker cycles 24/7. VDI delivers access to business information that is seamless for the worker enabling work anywhere/anytime in our 24/7 business world.

So how does deduplication impact all of these you might ask? Let’s take a look:
  • Data growth is all about scale and efficiency. We are saving huge amounts of data, but can we afford it? With data deduplication it can be reduced by 5-35X making it very affordable considering the alternative.
  • Transforming storage technology is about the fact that storage has not kept up with the advances in processors creating a performance gap that slows down overall worker and business efficiency. Solid state devices (SSDs) have a tremendous performance edge over existing spinning disk technology. However, they are hampered in their adoption because their price on a per GB basis is still too high compared to spinning disk. Deduplication closes that gap because it reduces the data stored on SSDs yielding a much better “effective cost.” In addition, deduplication addresses the other Achilles heel of SSD’s in that the less data written to the device, the longer the devices life cycle becomes.
  • Virtualization’s promise of efficiency can be optimized by deduplication. From a virtual server perspective dedupe of the server images can save both in the storage of the images, but also in the network utilization. From a storage virtualization view with deduplication data need only be deduped one time. As data’s importance evolves from hot to cool to cold, data is moved from expensive (SSD) to less costly hybrid storage and eventually to HDD and possibly tape. With the right deduplication technology, it never needs to be rehydrated between storage venues across the entire virtual storage environment. That will establish a virtual storage pool that will be more efficient (smaller) and more responsive to users as well reducing OPEX.
  • Virtual desktop infrastructure (VDI) basically creates a virtual “work” desktop that can be accessed from any device. The user sees their applications, and has access to their data no matter what portal they use. Deduplication optimizes that VDI creation at the server and the network access to the portal device. Dedupe saves the storage at the server and the network consumption by 20-35X in many cases.
Adoption of deduplication is a necessity as we transform information technology. Without it we cannot afford the data, access the trends that enable business differentiation or deliver access to data seamlessly on any device in the office, on the road or in your home.

There is one “gotcha” however, not just any deduplication can do this. Dedupe initially designed for backup isn’t even close, it was built just for backup. Deduplication that is needed to do this job must perform at I/O speeds and scale to multiple petabytes that first generation dedupe cannot meet. The dedupe engine that is needed to be broadly adopted across the IT landscape needs to deliver “performance at scale with efficiency.” Simply said, this next-gen dedupe engine must deliver:
  • Performance that meets today’s I/O requirements and does not incur latency 
  • Scalability that can scale to meet today’s high data growth environments 
  • Resource efficiency (RAM and Processors) which is critical to enable scale and performance 
All three of these - performance, scalability and resource efficiency - are needed to deploy a deduplication engine that can address the needs of business today and into the future. With increased data efficiency that deduplication can deliver, IT efficiency will improve markedly and an IT transformation will occur because costs will be reduced, data growth will be more affordable and will be in line with today’s IT budget constraints.

Tuesday, October 30, 2012

The Right Mix of Hosting, Storage and Security

- Brian Stillwell, senior director, solutions management, IT outsourcing, at Savvis, a CenturyLink company (www.savvis.com), says:


Posadas operates more than 102 hotel properties and 17,000 guestrooms in more than 50 different beach and urban destinations. The Mexico-based company provides a wide range of services to ensure its guests are met with fast and friendly service and its operations run smoothly.

It operates five unique hotel brands, including Live Aqua, Fiesta Americana Grand, Fiesta Americana, Fiesta Inn, and Hoteles One.

Challenge:
Posadas needed a more flexible, reliable and secure IT platform to help it manage a corresponding growing amount of data. As the business grows, the size of the information within its in-house data center also expands. The IT group found the costs required to manage the data becoming prohibitive. Meanwhile, the company expressed interest in implementing disaster recovery initiatives to recover lost data.

Posadas desired to outsource these IT requirements in order to avoid taking on additional permanent IT staff. The company also wanted to gain cost savings through the economies of scale that come with choosing a sophisticated outsourcer that would be able to combine multiple services.

Solution:
Posadas chose to work with Savvis, a CenturyLink company and an IT and could infrastructure leader. The hotel operator chose a full range of services, including managed hosting, storage, and dedicated servers. They chose Savvis for its ability to provide secure and available cloud-based hosting and responsive customer service. Savvis also partnered with Burstrom, a cloud infrastructure management firm, to develop an optimal cloud infrastructure implementation plan.

A core component of the infrastructure was Savvis Symphony Dedicated, a fully managed private cloud service. This solution allows enterprises such as Posadas to quickly scale its compute and storage resources in order to meet demand. Symphony Dedicated gives companies multiple public and private connection points while reducing the risks of failures.

Posadas also selected Savvis’ enterprise storage solution, which provides an accessible and scalable storage and disaster recovery platform that reduces the cost of ownership for enterprises of any kind. The solution reduces IT and management team and salaries while providing instant accessibility for any authorized user who needs to review, process or collect information.

Results:
Savvis’ managed hosting solutions allow Posadas’ staff to address its operations and improve company efficiencies instead of worrying about underlying infrastructure. Savvis’ managed hosting is an infrastructure-as-as-Service that features a dedicated service team to aid Posadas’ staff with any operational questions.

Savvis was able to quickly implement the Symphony Dedicated private cloud, seamlessly moving all of Posadas’ data center information without error. Savvis worked closely with Burstrom and Posadas to create an implementable cloud strategy in less than 120 days.

In order to protect against disaster, Posadas also chose enterprise storage from Savvis, providing the company with a safe repository for a massive amount of hotel-specific data, including information about all 102 properties and guest and marketing lists.

“We needed a cloud infrastructure provider that would be able to give us the right mix of hosting, storage and security at a reasonable cost to allow us to continue on our growth path,” said Leopoldo Toro Bala, CIO at Posadas. “Moving to the cloud is the right choice for any enterprise that wants to stay competitive, and by choosing Savvis we have become more agile and have the peace of mind of knowing our data will always be available.”


Monday, August 20, 2012

Recoup Storage Capacity Through Data Mapping

Jim McGann, VP of Information Discovery, Index Engines (www.indexengines.com), says: 


Understanding and profiling enterprise user data, its’ location and who owns it has been next to impossible for most organizations.  Regulatory, legal and compliance requirements are forcing organizations to better understand their data assets, manage it, and clean it up in order to avoid future liabilities. 

Data mapping can provide you the tools to better understand your data environment. Original data maps were created with face to face interviews, defining what data resides on which server and who owns it.  They don’t get to the granular level that is required by today’s demanding legal and risk management requirements.  Newer actionable data maps, based on enterprise class indexing technology, can help control costs and reduce resources, manage risk and liability and classify and manage data based on policy and storage rules. 

High speed, enterprise class indexing technology creates an actionable data map, providing comprehensive knowledge of data assets and a profile of user content to enable you to take further action.  High speed, efficient indexing will support petabytes of data which can be processed in reasonable amounts of time and with minimal resources.  This indexing platform must also support all classes of storage environments and data sources, including LAN filers, email databases, desktops and even backup tapes in order to ensure comprehensive knowledge of the enterprise.  This tool must allow for automated processes so that IT organizations can take action, such as migrating data to a legal hold archive, move to cheaper storage or defensibly delete what is no longer required.

Some use cases of a data map would be:
·         Find user data owned by ex-employees and determine if it has value to the business by reporting on last accessed data.  Purge data that has no value and reclaim storage capacity.
·         Find and manage all user pst’s email files.  Determine if it has been accessed or modified in the past 3 years, purge pst’s that are no longer in use.
·         Migrate user data that has not been accessed in 5 or more years to less expensive storage, including the cloud.

Data maps provide significant value to any organization and allow IT organizations to not only support legal and compliance with the information they require, but also to recoup storage capacity.  Analysts estimate that anywhere from 40 to 60% of storage capacity contains data that has no value or use to the organization and can be purged.  Implementing a data map is a cost effective strategy that will become a core component in tomorrow’s data center.


About the Authors
Jim McGann is VP of Marketing for Index Engines (www.indexengines.com), a leading electronic discovery provider based in New Jersey. McGann is a frequent writer and speaker on the topics of big data, backup tape remediation, electronic discovery and records management.

Thursday, July 19, 2012

Rebalancing Your Data Center Storage Solution


- Rob Commins, vice president of marketing, Tegile Systems (www.tegile.com), says:  

As we all know, application sprawl and data is growing exponentially because virtualization is enabling us to create and deploy applications faster and store data easier than any time in history. As we generate more data, we seek to preserve and protect our data with backup and replication, driving the demand for storage media even higher. The result is a significant challenge for IT departments, especially for those who want to consolidate IT infrastructure with cloud-based applications, virtualization and file sharing.

In order to maintain an efficient IT infrastructure that supports critical business operations and applications as virtualization becomes more pervasive, IT managers should look to rebalance their storage solutions. The means by which storage can be rebalanced includes: capacity, performance, compatibility, usability (fit for purpose), reliability, data protection and value for money.

Suggested action items necessary for rebalancing:

Capacity
Look to store more data per unit of rack space. With today's faster hybrid storage arrays more capacity can be packed in less than half the size of typical storage incumbents. Space in the data center is reduced along with power consumption.

Performance
Performance is always key in any discussion about storage. As performance increases, data processing time can be reduced from hours to just minutes - thus organizations need fewer servers, hard disk drives and fewer software licenses. The result of improved performance ultimately results in better services to an organization's IT users.

Compatibility
In order to create a truly unified storage environment, arrays should easily integrate into existing enterprise storage environments without server-side agents so they can work alongside or replace incumbent arrays. 

Usability
Make it easy on yourself. Make sure you can optimize virtual machines with just a few clicks so you can easily deploy many hypervisors or shares in minutes, not hours. Opt for systems that have graphs and customizable monitoring worksheets that make it easy to identify trends and issues for better planning and efficient optimization.

Reliability
This is such an important consideration. Look for systems that have no single point of failure architecture that includes dual hot-swappable controllers, dual power supplies and hot disk spares. Ensure that data is permanently stored on hard disk drives rather than flash drives that can wear out quickly in enterprise environments. That way, you'll enjoy a high level of resiliency to prevent both data loss and downtime without sacrificing performance.

Data Protection
Look to keep your applications online to improve recovery in the event of a problem. A great feature to consider is automatic snapshots and remote replication so critical machines can be backed up more frequently - that saves space and improves performance. And having the ability to roll back one machine or all machines to a previous state is a great feature. Ideally, only data that's changed should be backed up - requiring less network bandwidth, hardware and administration.

Value for Money
Look for arrays with built-in data reduction technology so usable capacity is greater than raw capacity and look for systems that don't require additional licenses for data backup features.

Following these guidelines will enable organizations that are budget sensitive with rapidly expanding storage infrastructures to stay ahead of the curve and maintain a high-performance storage model with all the functionality the modern data center needs.

Monday, July 9, 2012

Recent Enterprise Server Trends

- John Hodge, a writer for RackMountPro, writes:

Business owners ask a lot of their IT infrastructure. As the economy continues to lag, many are looking to do more with less by getting what use they can out of their existing systems.

At the same time, most businesses recognize the need to be both flexible and current to remain competitive. Today’s server challenges include making the workforce more mobile and getting the most performance for the money. Here are a few recent enterprise server trends that are designed to address these issues head on:

Petabyte Is the New Standard

If one gigabyte can hold seven minutes of high definition video, one petabyte can hold the equivalent of more than 13 years of high definition photo. Petabyte servers are, to put it simply, huge. They are also the new standard in enterprise storage.

Today’s businesses are making it standard practice to generate, gather, and store as much data as possible about their customers, their competitors, and their products which is later used to analyze industry trends. Petabyte servers give businesses and corporations the capacity they need to store this large amount of data without the need to install and maintain extremely large numbers of servers.

Improved Thermal Performance

Increased server density and the need for higher performance levels from existing servers have led IT professionals and server manufacturers to search for improved heat dissipation methods. Most of today’s enterprise servers are built with smart sensors that automatically control server thermals for optimized cooling.

These auto-control sensors also have the effect of saving energy, and thus, lowering energy expenses for the businesses that implement them. Some businesses even have the potential to increase their data center capacity by more than three times without significantly increasing energy costs or sacrificing performance.

Convergence for Simpler Management

Many businesses are seeking for simpler server management by converging their IT infrastructure. IT professionals are unifying software, storage, networking, applications, and more in order to increase synergy among the individual components of their data centers. This allows businesses to accelerate growth while containing costs.

Virtualization at All Levels

Although desktop and server virtualization is not a new concept, more and more businesses are “going virtual” from the ground up so that their storage, servers, and networking resources can be implemented throughout the organization. Virtualization allows employees to be more mobile—they can use their smart devices on the go or their personal laptops at home without sacrificing access to resources and business applications.

Virtualization also allows for increased security of sensitive data since it is all stored in a central location. This technique also lets IT managers to have more control of when and how information is used, and who has access to what files and programs. It also simplifies the IT management, which allows employees to spend more time working on mission-critical tasks and less time on mundane updates.

Storage Is Increasingly Non-local

Of course, perhaps the biggest trend in enterprise servers is cloud storage. About 40 percent of businesses currently use cloud services as either a primary or secondary backup for data, according to a recent survey. This number is only expected to rise as more businesses become more comfortable with the idea of outsourcing their storage needs.

The trend toward cloud storage affects nearly every level of today’s IT departments. What businesses gain in convenience and mobility, they sometimes sacrifice for complete data control and security. Whether businesses will jump fully “into the cloud” and give up their private data centers altogether remains to be seen.

These are only a few of the enterprise trends. What would you say are the most important trends in enterprise servers today?

John Hodge is a writer for RackMountPro. When he’s not writing he loves computers and everything related to them, gaming and spending time with his family. Connect with John on Google +. In addition to selling Windows and servers for Linux RackMountPro has been producing and selling rackmount servers and storage since 2001.


http://www.rackmountpro.com/

Monday, March 19, 2012

The Ins and Outs of IOPS

- Brad Bonn, senior systems engineer at VKernel (www.vkernel.com), says:

Recently, and over the last couple years or so specifically, I've seen IOPS become a buzzword everywhere. Phrases such as “We have a one million IOPS-capable SAN so storage shouldn't be having a problem,” or “I really need IOPS visibility” are popping up in conversations and appearing in online communities like the weeds in my garden. It makes sense that we're seeing the topic appear, because there is a decided lack of measurement standards for the overall performance of an end-to-end storage solution.

Earlier on in my years of IT, disk I/O capacity has classically been “measured” in terms of the number of acronyms you could rattle off. “Yeah, I've got a 12-spindle RAID 0+1 of 15K 146GB SAS drives in the backplane connected to my HBA with FC.” That's all well and good, but what does it translate to in terms of actual usability? How many databases can it support, and of what level of transactional utilization? How many users on an exchange system could it handle? What kind of file server load could it deliver? The universal answer is “it depends.” Application and file system diversity, the sharing of storage hardware through SAN/NAS devices, and the additional levels of sharing added through virtualization make unexpected results...well, expected.

The market abhors a vacuum, and when there is a clear need, vendors and integrators alike will move to try and fill it. The need in this case is the simplification of storage utilization both in terms of need from the application point of view, and in terms of delivery from the storage vendor point of view. Voila, IOPS.

It's a very simple concept, which is part of the reason it's become so widely used. “I/O Operations Per Second” is an easily understood and communicated unit of measurement. Unfortunately, it's also very easy to over-simplify. IOPS (or IOps or IOPs depending on the phrase you’re actually abbreviating) only describes the number of times an application, OS, or VM is reading and/or writing to storage each second. This sounds like a useful metric, because it is! More IOPS means more disk I/O, and if all IOPS are created equal we can measure disk activity with it alone. But the problem is, they aren't.

This topic is hotly debated on all sides, and having spoken with storage vendors, IT admins, and SMEs, I’ve come to the conclusion that as important as IOPS are, they aren’t the only metric you need to examine when you measure storage performance.

The goal of this document is primarily to outline the strengths and weaknesses of the use of IOPS in measuring storage capabilities; specifically from the perspective of shared storage. Along the way, we will cover some of the basic concepts surrounding shared storage itself and the implications that design choices in building a solution can have upon the performance and price of your infrastructure.

Shared Storage Fundamentals
If you’re already familiar with shared storage technology in general, feel free to skip ahead to the next section, but it’s helpful to review the components of a SAN to best understand the impact IOPS can have.

Where we keep the bits and bytes of data in our server farms has come a long way. Just spend some time on Wikipedia looking up things like “core memory” and punch cards to get a reminder of the evolution that storage has undergone over the years. Plus, not only has the medium by which we store data changed, but also the method by which we get information in and out of those sources has become just as diverse.

In the mainframe days, “shared storage” was a redundant title. All the computing resources for the building or company, including storage, were centrally located and therefore shared. Whatever tapes, disks or memory housed the data was all connected to the same core system, or systems. The resulting hierarchy was then logically very star-shaped with terminals connecting centrally in order to utilize the mainframe.

Modern computing is much more dense, and likewise, much more distributed. The giant mainframes of the past with their singular presences have been replaced by sprawling, interconnected datacenters consisting of dozens, hundreds, or even thousands of individual servers that all communicate internally and externally with other servers or personal computers. With each node being able to house its own storage devices, where data is located becomes equally as distributed.

These days, unless your “server room” consists of a few PCs and a SOHO router, you’re probably using a SAN or NAS in your infrastructure to share data between servers. Storage solutions like these make the allocation and migration of logical disks more manageable and fluid, to the point where directly-attached storage, or DAS (local disks contained inside each individual server) is practically never used in the datacenter any longer. Small infrastructures can still benefit from the lower initial investment of DAS, but in a rapidly-growing environment shared storage is critical to enabling rapid expansion, and vastly improving cost density. It seems the concept of where data is housed has come full circle and become centralized once more (failover sites notwithstanding.)

Whether you are using a storage area network or network-attached storage system, the end goal is the same: having a one-to-many relationship between a storage device and the computers that access it. Making that relationship happen involves various pieces of physical equipment and layers of logical abstraction, and while this whitepaper makes no claim to be a definitive guide on the broad topic of shared storage, we need to discuss some of the complexities involved in order to get a clear idea of what IOPS really means for such a system.

Links in the chain
Shared storage devices, regardless of their make, model, size or configuration all consist of the same general components. Any single read or write command has to traverse at least part of this chain in order to reach its destination, whether it is the active memory of a server, or the bare metal of a spinning disk. Every single link in this chain will affect the speed at which the data reaches where it’s going, as well as how many of those operations can be executed within a span of time.

The “weakest link” in this chain will determine the maximum number of IOPS and the total I/O bandwidth of the system. Every step of the way, additional connections are possible. In a block-level shared storage system such as fibre channel or iSCSI, (generally what would be considered a SAN) a LUN (logical unit number) can be accessed by multiple hosts (the effective one-to-many use case itself), but then those hosts can talk to multiple storage devices containing multiple disks via multiple paths.

In a file-based shared storage system such as an NFS filer (considered a type of NAS,) a LUN isn't used. Instead, filers will present network share locations, which can be accessed as logical disks by hosts and VMs. This option tends to be much more affordable since it can leverage existing networks and does not require as much specialized hardware. However this comes with a performance cost which may not always suit the needs of high-demand tier1 applications.

Starting at the “bottom,” we find the physical disks themselves. The faster the drive and the more data it can hold, the more expensive it becomes. Choosing the underlying disk technology will heavily impact the SAN or NAS’ total I/O capability and storage capacity. Lots can be done to get the most out of the spindles, but the buck stops here in the end.

A set of inexpensive magnetic SATA drives spinning at only 7200RPM bring a great cost to storage density ratio, but at the expense of limited speed. I/O-light applications or data archives are always well-suited to these. On the other hand, a set of 15,000RPM serial-attached SCSI (SAS) drives with hefty memory caches on-board, or solid-state drives (SSD) with no moving parts built entirely from flash chips will bring astonishing speed to the table for hefty database operations or disk-intensive apps at the cost of serious hit to your budget. Larger SANs will contain a mixture of these kinds of drive technologies, allowing for the intelligent balancing of disk loads to match performance tiers.

On the next link in the chain, these physical disks are bundled into logical groups, or arrays, often known as a RAID (Redundant Array of Independent Disks) which ensure the protection of the data against disk failure, and can potentially speed up bare disk operations through parallel reads and writes. However, depending on the type of RAID configuration in place, the total performance of the disks can be negatively impacted as well. For a detailed description of the types of RAID that are out there and their effects on performance, this article online can help: http://www.accs.com/p_and_p/RAID/BasicRAID.html

Many modern SANs take array configuration completely out of the hands of the storage administrator. This can greatly simplify the bare disk component of the storage and enforce best practices across tiers.

In order for the disks in an array to function as a logical unit, there must be a device or software to organize them into such a structure. In a shared storage configuration, the storage processor (or storage controller) handles this. Usually housed inside the chassis of the SAN, the storage processor is a self-contained computer system that handles all I/O for the device. It manages each I/O operation written to and read from the disks, as well as manages the communication to the hosts utilizing the shared storage via the various mediums available.

A storage system can have multiple redundant or load-balanced storage processors, and can be made up of multiple physical chassis containing anywhere from a few dozen to a few hundred physical disks.

With both SAN and NAS shared storage models involving a many-to-many connection scheme between the physical drives and logical disks, the full map of a complex SAN configuration can become quite the atlas!

Each of these links in the chain will affect the end-to-end performance of a SAN or NAS and the diverse possibilities that lie in each possible portion mean that shared storage configurations can vary to an astonishing degree.

Troubleshooting poor performance in a shared storage system can often feel like trying to search the inside of an oil tanker with a penlight. And while the specifics of what is causing slowness can be innumerably varied, the source or sources of storage limitations typically break down to a few major areas. These are where storage admins generally look:

• Hardware or system failures
• Communication “traffic jams” to the storage device
• Surges in storage I/O load demand
• Inefficient configurations for required throughput
• Application or OS-level inefficiencies

Failures in the chain of storage will cause at worst a complete outage, or at the very least a disruption in the performance of the complete system. For example, disk failures will place a RAID array into a “degraded” state where the system attempts to work around the missing disk or disks. Degraded arrays will almost always be significantly slower in responding to IOPS and degraded arrays no longer have their redundancy and could exhibit data loss if further failures occur, so be sure to replace the failed disks ASAP. In addition to the slowdown caused by the degraded array, once a replacement disk is inserted, the array must be rebuilt. This causes the storage processor and all disks in said array to have an additional task to perform on top of whatever normal I/O load is being placed on it. This will slow down the responsiveness of the system even further.

There can also be failures in the communication between the hosts and the storage device. Loss of connectivity on redundant links will cause a reduction in the maximum amount of bandwidth to the SAN/NAS, or even sever connections entirely, causing outages or failover.

The storage devices themselves can also experience outages and failures, whether it is in the form of a storage controller failure, an issue on the backplane circuitry, or any other hosts of potential breakdowns.

Generally, any good storage system has redundancy throughout, allowing for various components to fail without bringing production operations to a standstill. For this reason on top of the fact that failures are rare, this tends to be the least likely cause of performance problems in a shared storage configuration.

Outside of system failures, excessive I/O loads on a perfectly functional storage system can create higher latencies and therefore cause poor application performance. While this is the most obvious potential issue, it is also one of the most difficult to diagnose. Closely monitoring I/O load metrics such as IOPS and MB/s are critical for determining where heavy loads are coming from and how to most appropriately respond to those loads.

Visibility into the various “links in the chain” to determine the source of these metrics is also critical in narrowing down which portion of the storage system is taxed most heavily. Is the disk array overloaded, or are the iSCSI links saturated? Is the storage processor’s CPU unable to keep up, or are the host bus adapters inside the servers being pushed to their limits? If iSCSI or another Ethernet-based connection is employed, are jumbo frames turned on at every step of the way? Very frequently, VM Admins don’t realize that even in vSphere 4.1, setting up multipath iSCSI with jumbo frames involves a lot of legwork in the ESX console before it works fully! This is being improved upon in 5.0 to be more GUI-centric, but double-check your work using the guide here to make sure:

http://www.vmware.com/pdf/vsphere4/r41/vsp_41_iscsi_san_cfg.pdf

Occasionally, the bottleneck might be within the application or OS, and not in the storage environment at all. For example, poor I/O performance on a database server might be due to tiny growth increments, which cause fragmentation. Another common (but decreasingly so) example is disk alignment. If the clusters of a partition are not aligned with the blocks of the disk (either physical or virtual) then reading one cluster may end up requiring access to up to three chunks of the LUN. Under extreme conditions, a misaligned disk can cause up to a 40% degradation in I/O performance! Most modern operating systems don’t exhibit this problem, and typically it only occurs in legacy deployments.

In the end, working out bottlenecks in the chain involves figuring out which component is waiting for the other to finish its job and send the information along. Any time the average total latencies of I/O operations exceed 20ms, the applications and the users on them will begin to notice degradation of performance. As this latency increases, performance will only get worse, potentially reaching the point where the delay in reads and writes will cause the applications or OS’s hosting them to time out and give up on their attempts to access the disk. In virtual and non-virtual systems, these situations are often logged as “aborted I/O commands,” and they can cause serious high-level errors if proper handling for I/O is not implemented within the applications.

Measuring the I/O latency is most often done from within the SAN. Tools from storage manufacturers or third-parties will grant visibility into the wait times for IOPS and can help pinpoint the source of the slowdown from within the storage system. However, these tools often overlook the guests themselves and their point-of-view. Measuring storage latency at the server (or VDI) is just as important, since latency inside the SAN doesn’t always account for the links and protocols which connect that SAN to all of the systems communicating with it. Looking deep into the storage device will help show which RAID array is having the most trouble, for example, but it will not help track down the fact that one of the network connections is saturated, causing the traffic to the SAN to bottleneck. So make sure that whatever methodology you are using to monitor disk performance incorporates a full view of latency end-to-end with the goal being to keep those numbers as small as possible.

What does all this mean? It means storage is complicated! Despite the simplicity of the concept to store a byte in a location and being able to access it from many places, the implementation of such becomes a mechanism of underappreciated complexity.

IOPS as a measurement of disk I/O
So, back to IOPS. Where does this measurement metric come into play, and how does it affect the overall picture of disk I/O? Can your storage system’s performance be measured in IOPS effectively? Well, yes, but only in part. A better question to ask is, “What do you define as performance?” Are we talking about the maximum I/O potential of a storage system, or the responsiveness of the storage to the demand being placed upon it?

IOPS are an effect of a storage system’s performance. Better performance on a more expensive SAN means more potential IOPS. Easy, right? Well here’s where things become tricky:

An idle storage system has zero IOPS. As load increases on a storage system, the IOPS against it go up. As the load continues to increase, bottlenecks within the system will cause latency to rise and with more time between each I/O, the number of IOPS will eventually plateau. IOPS can therefore be an indicator of load under ideal circumstances, but one cannot simply say that because a SAN is showing a certain number of IOPS that it is performing well. It’s not until the latency for those IOPS is examined that we know the SAN has reached its saturation point and our applications have begun to suffer.

Storage performance affecting IOPS
There are many, many ways to build a storage system. And nearly all of these permutations will have an affect upon the number of simultaneous inputs and outputs that can be executed.

For example, let’s say a single spindle can perform X number of IOPS. However, if that disk is now striped with a second disk inside an array, this increases the amount of possible I/O by giving the potential to read and write to both devices simultaneously. Thereby giving us X * 2 total IOPS available. Correct? Well, sort of. Even though the I/O is going to two disks at once, thereby increasing the total amount of data that can be written simultaneously, the application layer does not see this. The storage processor is transparently handling the transfer to both disks, and with the increase in bandwidth, can potentially handle more IOPS, but it isn’t as simple as pure multiplication.

The process of encapsulating I/O also involves overhead on the part of the storage processor itself, so while adding more spindles to an array will bring about more available I/O capacity to a point, eventually the overhead becomes so great that all benefit of the parallelization is lost. This concept holds true throughout computing, so I won’t address it in-depth in this article. For additional reading on the concept, point your browsers here:

http://en.wikipedia.org/wiki/Parallel_computing

The potential bandwidth of a storage system also affects IOPS, regardless of an I/O’s size or nature. The more bits that can be sent per second, the more IOPS, so it’s clear that a SAN’s performance affects IOPS.

IOPS affecting storage performance
Now let’s look at the situation from the other side. Reading or writing to a storage device means that those IOPS are placing load upon the network links, taking up CPU cycles on the storage processors and HBAs, and making the heads of the spindles move to various locations around the disks. This means that if another system wants to utilize that same shared storage device, it will need to work within the boundaries of what remains. The disk heads will probably now have further to travel in order to read or write the next time around and the network links only have so much bandwidth remaining at that moment, etc. This means that the more IOPS that are being pushed through the storage, the slower the response time will be for other systems trying to access it. Even if one system has a relatively low demand on the disk, another with high demand will cause that low-demand one to have slower performance waiting with its foot tapping for the I/O request it sent to come back from the SAN.

This means that the number of IOPS is both an indicator of the SAN’s load and an affecter of its “performance.” It really depends upon the perspective we are examining in the situation.
Not all IOPS are created equal until they enter the storage system
The metric of I/O’s per second is one that involves several caveats alluded to earlier in this paper. What size are the I/Os themselves? What percentage of them consists of read operations vs. write operations? Are they seeking data that’s likely to be read from sequential areas of the storage or are they utterly random?


This is a screen capture from just one environment where the disconnect between IOPS and total I/O load is markedly visible.

Notice the fact that the first VM is showing only 22% more IOPS than the second VM, yet it’s showing 800% more disk throughput! Looking at throughput shows that the first VM is vastly heavier in disk I/O, yet if we were only considering IOPS, these two VMs on the same datastore would be almost indistinguishable between each other in their I/O needs. If these VMs have high disk latency (waiting long periods for their I/Os to return) how would you then determine which VM should be granted a more dedicated or higher-performance LUN for its operations in order to reduce that latency? Purely based on IOPS, it would practically be a coin toss. But since one VM is writing 64KB blocks and the other appears to be writing 12KB blocks, the difference is much more drastic.

Just looking at block sizes alone, a storage system’s I/O capacity will vary greatly even with the same number of IOPS. In short, the bigger the IOP, the fewer of them that can go through the entirety of a storage system.

Conversely, the throughput potential of the entire system actually increases with larger block sizes. Much like enabling jumbo frames on an Ethernet connection can decrease overhead and improve maximum throughput for larger data transfers, the same can be true for disk I/O in certain circumstances.

Note that the total amount of I/O data throughput (essentially, the storage bandwidth) of a storage system plateaus at a certain block size and does not continue to improve. This is a perfect example of how the principle limitation of a storage system can be the total amount of bytes that can be written to or read from it each second and not so much the number of times a read or write operation can be executed on that system per second.

It’s important to note that this differentiation ends once an IOp enters the physical storage device itself. Once inside the storage processor, the block sizes are consistent. However, when examining the end-to-end performance of a storage solution, this differentiation is vital.

So when trying to design and configure storage, do you focus on maximum potential I/O’s or maximum throughput? The answer will depend on the following factors:

1) The class of operation
a. Higher performance expectations and SLA’s generally will demand low latency numbers, so storage must be not only optimal, but capable of sustained operation at high loads.
2) The type of data being accessed
a. Large quantities of tiny files or granular database operations will benefit from systems that perform better with small block sizes, yielding greater IOPS for when throughput is secondary.
b. Larger files or data access taking place in bigger chunks won’t take advantage of smaller block sizes, and therefore will not benefit from an IOPS-oriented design, instead needing top performance for sustained throughput.

Let's take a VDI infrastructure for example:

In most virtual desktop deployments, disk I/O tends to be very heavy during initial boot-storms with close to 99% of the IOPS being read from the boot image(s) that are very similar to one another. (Powering on a bunch of virtual desktops that all run the same operating system at the beginning of a work day.) Then, during normal operation the majority of IOPS are frequent, small, but very random writes alongside random reads. (Periodic saving of work files, writes to web browsing cache folders, email client activity when receiving messages, etc.) This kind of activity benefits heavily from caching whether on the spindles themselves, in the storage controller, or the HBA. During the boot periods all of the reads are from very similar or identical images, so caching means only one read IO from the spindles is needed for each sector, leaving the bottleneck to be the maximum throughput speed of the fabric to the SAN, provided the total caching size is large enough to store the full “golden” boot image within memory.

Once booted up, the virtual desktops issuing write commands will also benefit from write caching, allowing extra time for the spindles to keep up with receiving the I/Os. A 4GB write-back cache could provide space for over one million 4KB-sized I/O blocks (a common block size,) giving ample time for them to be written to disk in between bursts of disk activity. Once again, the bottleneck would usually come down to the media fabric between the storage and the hosts. So in the case of VDI (often touted as a very IOPS-intensive function,) a SAN with ample caching would support a great deal of VDI-oriented IOPS before suffering any kind of slowdown either from the cache filling up, or from the HBAs being unable to deliver data to the SAN fast enough.

Storage architects will quickly (and correctly) note that any sort of caching will only improve I/O performance at burst speeds, and primarily for write operations. Read operations only occasionally benefit from caching because they depend on the data already existing in the cache’s memory, either from being read previously or from an algorithmic selection of data that is likely to be read in the near future. (The percentage of success in utilizing a cache resource is known as its “hit rate”.) So while this might work ideally in the case of VDI, the same configuration may not do nearly as much good for a large-scale database deployment

Making the Most of IOPS
With all of the complexities involved with storage, the divergent nature of throughput against IOPS, and the fact that no two IOPS are the same, does this mean that IOPS is a worthless metric? Far from it! It’s merely one of many pieces of the puzzle of disk I/O. It also means that consistency is key, and that all the factors involved need to be accounted for in order to determine the performance requirements of an infrastructure, as well as the capabilities of a storage system to handle those requirements.

Consistency needs to come from the perspective of the storage vendor, so that when a system architect is looking to choose a storage solution, they have an even playing field. If a vendor promises a system capable of one million IOPS, and all of those IOPS are sequential, read-only with 100% cache success, and bursting for no longer than 10 seconds, then that information is at best, not helpful, and at worst, false advertising. An industry standard for overall I/O measurement is something that I believe the community should call for, but in the meantime, following these guidelines can help keep things in perspective when choosing a shared storage solution:

1) Assume the worst-case scenario and avoid over-simplification

Make sure that the storage vendor provides detailed information about the number, size, and type of IOPS the system can handle. Storage manufacturers tend to optimize their hardware and software for 512 byte transfers, to maximize their advertised IOPS rate. But if IOPS aren’t as important as throughput for your application, this won’t be optimal. Plus, 4KB or 8KB transfers are far more realistic to encounter in real-world applications.

Ensure that the numbers assume a 0% cache hit rate, eliminating potential false performance improvements due to rare “ideal” circumstances.

Arm yourself with information from application vendors ahead of time to work out what kind of I/O requirements you’re going to be expected to support within the infrastructure. Lots of tiny I/Os vs. a small number of very large transfers will place very different requirements on your storage.

2) Hold each solution to the same standard

Even if you’re unsure of the exact I/O needs of your application, use the same measuring stick between vendors and models.

Keep block sizes, stripe sizes, drive technologies and spindle counts consistent to narrow down the focus to the storage processor(s) and interfaces. Very frequently, Vendor A will use the exact same disks inside the SAN as Vendor B. Your focus, therefore, should be on working out how well their system can handle those disks.

3) Diversify, Diversify, Diversify

Don’t put all your eggs in one basket. If you have a complex environment with variable storage performance needs, you should not be held to just one storage solution or model. Even if you keep a consistent vendor, focus on application-level delivery when determining the storage configuration that will work best for you.

And even if you want to keep a consistent vendor AND model, diversify the SAN itself! Nearly all storage systems have provisions for more than one type of disk, storage processor, and communication medium.

Once a storage system is in place, following best practices in allocating space within it is critical to getting the most out of the deployment. For instance, when configuring LUNs it is important to divide and conquer. Do not put all of your disks into one logical array and expect great performance! Depending on the number of disks, the parity calculations alone could push your storage processors to 100% even with the lightest of I/O loads.

It is better to split things up. Put your I/O-light apps on inexpensive disks in parity configurations, and dedicate high-performance drives in their own LUNs for databases or other I/O–intensive operations. This can almost always be accomplished inside the same physical SAN.

Storage vendors will have their own technologies and methods for assigning the right amount of disks in the right configuration for best performance and most efficient distribution of available space. Make sure to work closely with the manufacturer whenever possible, since even someone with a great deal of storage experience can be taken by surprise.

4) Check your (and their) work
Verify that the promised amount of I/O capacity lines up with what should be theoretically possible given the hardware configuration proposed. Calculators like the one here by Marek WoƂynko can help: http://www.wmarow.com/storage/strcalc.html

Utilize tools such as I/O Meter: http://www.iometer.org/ to push storage systems to their limits and see how well they perform under high-stress loads before they end up being put into production.

Once in production, closely monitor I/O responsiveness from the perspective of the systems utilizing the storage. Keep an eye on your latency values. No IO operation should be taking longer than 50ms to round-trip within the storage environment in order to maintain the best performance and even tighter tolerances will be needed for higher-tier operations.

Conclusion
It is clear that IOPS as a term for storage performance is here to stay, regardless of any limitations that may exist with its use. The best way to keep on top of what it means to you and to your infrastructure is to remain informed. This whitepaper began as a blog post just about IOPS, but this topic is so enormous that I am going to keep this as a “work in progress” that I will keep updating with time.