Since the need for a reliable data archive is well known, there’s no need for us to focus on that. Instead, we’ll discuss the various data archive options AWS offers its customers. However, we should first make an important distinction between data archives and data backups – as the purpose and function of the two should not be confused.
A data archive is for data not actively in use, but that needs to be moved to a separate storage device for preservation and retention over the long term. Besides preservation, a key goal of a data archive is to reduce the cost of storage. A data archive is not intended to help your system recover from some disaster or failure. Backups – which are performed on both active and inactive data – are designed to permit recovery from data failure.
Bearing that in mind, being able to quickly restore data from a backup medium is likely far more important than it would be for an archive. Such considerations will define the kind of ideal solutions you might choose for your data archive vs. your data backup.
The data archive: traditional considerations
- You would need to see far into the future, as the medium you are using today may not exist in ten years. It can therefore be a real challenge to identify a viable long-term storage platform.
- While archived data may not be currently active, they are generally intended for production use. Therefore, reliable security over long periods of time becomes a critical goal.
- Data archives tend to grow with time, so you will need to realistically consider future costs and scalable infrastructure needs upfront.
- Most organizations – and especially governments – are very particular about the availability of archived data. The process of meeting such expectations may lead you to improved disaster recovery strategies, but it can also be really complicated and expensive.
- Implementation can require significant skills and experience across multiple technologies.
However, many of these concerns simply wouldn’t apply to a data archive in the Cloud. Using AWS, for instance, means you never need to invest in a particular technology or medium, or worry about changing standards – that’s all Amazon’s headache. And your costs will always be a direct product of the services you actually use.
Your data archive and the AWS Cloud
Of course, AWS isn’t the only player offering out-of-the-box cloud archiving services, but they’re a good place to start.
S3 and Glacier are, one way or another, the primary AWS tools you’ll use for your data archive. We’ll look at three common use-case scenarios: archives for AWS-based data, on-premise data, and hybrid data solutions.
1. Applications deployed within AWS
If the application to be archived is running within the AWS environment, then integration with S3 or Glacier should be straightforward. Since a data archive doesn’t demand frequent reads, you would normally opt for the cheaper Glacier, which can require a lag of several hours for retrieval. If, however, you’re already storing some application data in S3 (like videos or application logs) and you may not want to write the extra code needed to move inactive data to Glacier, you may instead consider moving only the old, inactive data from S3 to Glacier.
AWS allows you to configure and manage the automated lifecycle of objects in your S3 buckets. You could therefore create a configuration that causes S3 objects to be moved to Glacier based on specified conditions or policies.
A sample policy may look like this:
2. Applications deployed on premises
If the components of your application (like a webserver, database, application server, and NFS server) are running within your datacenter, but you still want to use AWS for archiving your backed up data, the simplest solution is to integrate your backup server with AWS S3 or Glacier. This diagram may help you visualize the architecture:
If you’re already using AWS S3 for your backups instead of a local backup server, then you can use S3 Lifecycle management to quickly add a data archive layer using Glacier to your infrastructure.
Even if your backup server doesn’t natively support AWS cloud integration, you can still create a seamless and secure interface between your data center and AWS’s storage infrastructure using AWS Storage Gateway. Storage Gateway won’t require a dedicated network setup between your corporate network and AWS infrastructure, and it is built to support industry standard storage protocols, while storing the encrypted data in AWS S3.
3. Applications deployed in a hybrid setup
In this kind of setup, an application deployed on AWS might interact with on-premise components (or the other way around). In such cases, you may want to extend an existing archiving strategy to the cloud, requiring only a reliable way to connect your two networks via either a standard VPN setup or through AWS Direct Connect, which makes it easy to establish a dedicated network connection from your premises to AWS.
Data archive compliance and regulations
Many customers will have specific data retention policies, and must often comply with regulatory guidelines. AWS Glacier offers you Vault Locks. A Vault Lock Policy allows you to apply compliance controls to the contents of any Glacier vault.
To review, here are some of the key advantages you can enjoy by archiving your data in the cloud…and with AWS in particular:
- No more need to rely on risky predictions of your data growth and corresponding data storage.
- Reduced overhead of managing huge data stores for long periods.
- Reduced cost.
- Increased availability.
- No more need to identify and invest in some particular hardware and skills to implement a reliable, long-term archival design.
Do you have your own cloud/local archiving experience? Let us know in the comments.
New Content: Platforms, Programming, and DevOps – Something for Everyone
This month our team of expert certification specialists released three new or updated learning paths, 16 courses, 13 hands-on labs, and four lab challenges! New content on Cloud Academy You can always visit our Content Roadmap to see what’s just released as well as what’s coming soon....
Mastering AWS Organizations Service Control Policies
Service Control Policies (SCPs) are IAM-like policies to manage permissions in AWS Organizations. SCPs restrict the actions allowed for accounts within the organization making each one of them compliant with your guidelines. SCPs are not meant to grant permissions; you should consider ...
New Content: Focus on DevOps and Programming Content this Month
This month our team of expert certification specialists released 12 new or updated learning paths, 15 courses, 25 hands-on labs, and four lab challenges! New content on Cloud Academy You can always visit our Content Roadmap to see what’s just released as well as what’s coming soon. Ja...
New Content: Get Ready for the CISM Cert Exam & Learn About Alibaba, Plus All the AWS, GCP, and Azure Courses You Know You Can Count On
This month our team of intrepid certification specialists released five learning paths, seven courses, 19 hands-on labs, and three lab challenges! One particularly interesting new learning path is Certified Information Security Manager (CISM) Foundations. After completing this learn...
Which Certifications Should I Get?
The old AWS slogan, “Cloud is the new normal” is indeed a reality today. Really, cloud has been the new normal for a while now and getting credentials has become an increasingly effective way to quickly showcase your abilities to recruiters and companies. With all that in mind, the s...
The 12 AWS Certifications: Which is Right for You and Your Team?
As companies increasingly shift workloads to the public cloud, cloud computing has moved from a nice-to-have to a core competency in the enterprise. This shift requires a new set of skills to design, deploy, and manage applications in cloud computing. As the market leader and most ma...
AWS Certified Solutions Architect Associate: A Study Guide
Want to take a really impactful step in your technical career? Explore the AWS Solutions Architect Associate certificate. Its new version (SAA-C02) was released on March 23, 2020. The AWS Solutions Architect - Associate Certification (or Sol Arch Associate for short) offers some ...
New Content: AWS Terraform, Java Programming Lab Challenges, Azure DP-900 & DP-300 Certification Exam Prep, Plus Plenty More Amazon, Google, Microsoft, and Big Data Courses
This month our Content Team continues building the catalog of courses for everyone learning about AWS, GCP, and Microsoft Azure. In addition, this month’s updates include several Java programming lab challenges and a couple of courses on big data. In total, we released five new learning...
Where Should You Be Focusing Your AWS Security Efforts?
Another day, another re:Invent session! This time I listened to Stephen Schmidt’s session, “AWS Security: Where we've been, where we're going.” Amongst covering the highlights of AWS security during 2020, a number of newly added AWS features/services were discussed, including: AWS Audit...
AWS re:Invent: 2020 Keynote Top Highlights and More
We’ve gotten through the first five days of the special all-virtual 2020 edition of AWS re:Invent. It’s always a really exciting time for practitioners in the field to see what features and services AWS has cooked up for the year ahead. This year’s conference is a marathon and not a...
WARNING: Great Cloud Content Ahead
At Cloud Academy, content is at the heart of what we do. We work with the world’s leading cloud and operations teams to develop video courses and learning paths that accelerate teams and drive digital transformation. First and foremost, we listen to our customers’ needs and we stay ahea...
Excelling in AWS, Azure, and Beyond – How Danut Prisacaru Prepares for the Future
Meet Danut Prisacaru. Danut has been a Software Architect for the past 10 years and has been involved in Software Engineering for 30 years. He’s passionate about software and learning, and jokes that coding is basically the only thing he can do well (!). We think his enthusiasm shines t...