Sunday, January 25, 2015

modern compression tools are fast !

This morning I played with some compression tools on a new 28 core machine (Xeon E5-2695 v3 @ 2.30GHz).

I used python to create a 1GB string that consists of fake random DNA and dumped that to a text file:

   import sys, random  
   dnalist= list('ACGTACGTACGTACGT')  
   bytesize=1024*1024*1024  
   hostname=socket.gethostname()  
   dnastr = ''   
   for i in range(bytesize):  
     dnastr += random.choice(dnalist)  
   sys.stdout.write(dnastr)  


then I tried standard gzip as well as the new lz4, lzo  and the highly parallel pigz compressor which produces gzip compatible archives:

Tool compression level file size (MB) run time (s)
gzip 6 293 111
lz4 1 693 15
lz4 6 466 69
lzo 6 500 6
lzo 7 399 412
pigz 6 292 5

lz4 performance is certainly an improvement over gzip at the price of lower compression ratio. However in this test it is not quite as impressive as in these benchmarks. https://code.google.com/p/lz4/

lzo is doing really well and is actually much faster than lz4 while delivering similar compression. Level 7-9 are really not that useful though.

lz4 claims to have much faster decompression times than lzo but I cannot confirm this here. Both tools take about 6 seconds to decompress and restore the 1GB file.

pigz shows what can be done with raw compute power. top showed 2800% cpu utilization on this 28 core linux system. It seems to scale almost linearly to the numbers of cores. Decompression takes about 3 seconds. Here the local raid array may be a limiting factor. It can write 300-400 MB/s




Monday, October 14, 2013

Linux Kernel Roadmap for Ubuntu LTS

Since we are mostly a Ubuntu Shop we are somewhat interested which Kernel will be in the next version 14.04 LTS. Since there is no official road map we have to do a little crystal balling. Let's see how long each release actually takes in the 3.x series. Taking the release dates from Wikipedia we see an average on 67 days in the last 2 years:

Version      Release Date           Days
3.0    7/22/2011    64
3.1   10/24/2011    94
3.2    1/05/2012    73
3.3    3/19/2012    74
3.4    5/21/2012    63
3.5    7/12/2012    52
3.6   10/01/2012    81
3.7   12/11/2012    71
3.8    2/19/2013    70
3.9    4/29/2013    69
3.10   6/30/2013    62
3.11   9/02/2013    64

and assuming that Linus will turn into a robot and releases every 67 days the future may look like this:

3.12   11/8/2013    67
3.13   1/14/2014    67
3.14   3/22/2014    67
3.15   5/28/2014    67
3.16    8/3/2014    67
3.17   10/9/2014    67
3.18   2/15/2014    67
3.19   2/20/2015    67
3.20   4/28/2015    67
3.21    7/4/2015    67
3.22    9/9/2015    67
3.23  11/15/2015    67
3.24   1/21/2016    67
3.25   3/28/2016    67

For Ubuntu 14.04 this means that the Kernel will either be 3.13 or 3.14, the former is perhaps more likely.


Thursday, March 28, 2013

"ZFS on Linux" ready for wide scale deployment


Quoting lead developer Brian Behlendorf from Lawrence Livermore National Lab (LLNL):

"Today the ZFS on Linux project reached an important milestone with the official 0.6.1 release! Over two years of use by real users has convinced us ZoL is ready for wide scale deployment on everything from desktops to super computers." Read the full announcement

ZFSOnLinux (or ZoL) is a high performance implementation of ZFS as a Kernel module and performance is on par with Solaris (especially on new Hardware).

It is used by the LLNL Sequoia HPC Cluster with 55PB of storage:

http://arstechnica.com/information-technology/2012/06/with-16-petaflops-and-1-6m-cores-doe-supercomputer-is-worlds-fastest/

The porting of ZFS to Linux has been funded by DOE and did start in 2008. Is it important to understand that ZoL is not currently an unstable beta but the result of a more than 5 year effort.

Please see presentations from 2011 and 2012 that provide additional details:
http://zfsonlinux.org/docs.html

Sunday, February 17, 2013

Ubuntu 12.04.2 LTS Kernel confusion

Recently there were some changes in Ubuntu:With the release of 12.04.2 it seems new installs from cd/dvd will use the lts-quantal kernel 3.5 from Ubuntu 12.10 by default..... however apt-get upgrade and dist-upgrade will continue to default to the old 3.2 kernel. The 3.5 kernel will enjoy the same support the 3.2 kernel had, but the 3.5 kernel will only be supported until the next LTS release 14.04 while the 3.2 kernel will be supported for the full 5 years. Canonical recommends to leave VMs and cloud installs at 3.2.
https://wiki.kubuntu.org/Kernel/LTSEnablementStack
Since we want to keep our Scientific Computing stack fresh this would mean that we upgrade to 14.04 next year. This puts one question on the table: Do we want to upgrade our current compute systems to kernel 3.5 or stay on 3.2?
Not yet sure if there are any direct benefits other than better support for the Micron SSD controller in the Dell R720 hardware we use. One that I could see is that it supports tcp connection repair which is useful for HPC checkpointing.
Another interesting feature is improved performance debugging.
http://kernelnewbies.org/Linux_3.5#head-95fccbb746226f6b9dfa4d1a48801f63e11688de
and a network priority cgroup:
http://kernelnewbies.org/Linux_3.3#head-f0a57845639c0fbc242438e4cb76d44d1f103c24
we would probably leave most of our virtual systems on kernel 3.2 to enjoy the full 5 year support, our desktop deployment should may be go to 3.5 if the hardware requires it.

Wednesday, August 29, 2012

OpenStack Swift vs AWS Glacier costs

Since AWS Glacier hit the road there are some interesting discussions and blogs on comparing costs between Glacier and local solutions such as OpenStack Swift. Glacier is hard to beat if you do TCO calculations that include everything like datacenter, power, cooling & staff. For many of us these costs vary a lot dependent on things like being in a fortune 500 or in a government funded agency or residing in a location with low power and cooling costs vs LA or NYC. Some of us even have the notion of sunk costs .....

If we just look at the plain storage hardware we know for example that we can get 36 drive standard Supermicro storage servers for less than $5k and we have seen the latest and greatest 4TB Hitachi Deskstar for $239 on Google Shopping. The Hitachi Deskstar model seems to have an excellent reputation and folks who know what they are doing recommend it as well. (albeit the older 3TB version).
So we seem to be getting 144TB RAW which might roughly translate to 130TiB usable in this box and it costs ($5000+36*$239)/130TiB = $105-$115/TiB dependent on your sales tax... let's say $110/TiB. Swift needs 2-3 replicas so your actual costs would end up at $330/TB or $66/TB/Y if we assume that the whole system will run for 5 years. That's not too bad compared to Glacier which runs minimally at $120/TB/Y.
If swift sounds compelling to you, you still have to operate and support it but you can actually get tech support from a number of vendors such as www.swiftstack.com .

Amar here has another idea which I find intriguing. LTFS allows you to mount each tape drive (up to 3TB capacity each) into an individual folder on your Linux box. Just using LTFS is probably painful since you may have hundreds of small 3TB storage buckets ......but if there was a way to use Swift with LTFS this could possibly push down storage costs to under $20/TB/Y. I'd like to learn more about this.





Sunday, May 13, 2012

OpenStack Swift vs Gluster

As I am trying to get my head around OpenStack Swift storage I need to compare this to something we already know. We have been using GlusterFS for years in our shop and are reasonably happy with it for data that does not require high performance disk and high uptime. Gluster sounds like a simple solution but its codebase has grown over the years and it has not been free of bugs. As of 2012 it is really quite stable.

Let's look at the 2 codebases:

git clone https://github.com/gluster/glusterfs.git
git clone https://github.com/openstack/swift.git

>du -h --summarize glusterfs/
44M     glusterfs/
>du -h --summarize swift/
15M     swift

Well, gluster is 3 times the size, let's take a more detailed look at the code:

>cloc --by-file-by-lang glusterfs/


---------------------------------------------------
Language         files       comment           code
---------------------------------------------------
C                  272         14179         256462
C/C++ Header       214          5289          23208
XML                 24             2           6544
Python              25          1836           5114
m4                   3            85           1447
Bourne Shell        34           359           1419
Java                 7           168            988
make               107            36            965
yacc                 1            15            468
Lisp                 2            59            124
vim script           1            49             89
lex                  1            15             64
---------------------------------------------------
SUM:               691         22092         296892
---------------------------------------------------


>cloc --by-file-by-lang swift/


---------------------------------------------------
Language         files       comment           code
---------------------------------------------------
Python             101          6137          32575
CSS                  3            59            627
Bourne Shell         8           138            251
HTML                 2             0             82
Bourne Again Shell   3             0             23
---------------------------------------------------
SUM:               117          6334          33558
---------------------------------------------------

Hm, gluster has 8 times more lines of C code (SLOC) than swift has python code. I'm not in the position to compare python with C (other than stating that as of 2012 they seem to be similarly popular) but if we simply assumed that the numbers of errors per lines of code is similar swift may at some point have a stability advantage over gluster. Gluster has been developed for many years and it took a long time to come along. Swift is only been around for 2 years and some really big shops seem to be betting on it. Of course this is somewhat an apples to oranges comparison because Gluster is accessible as posix file system and object store and also has it's own protocol stack (NFS/glusterfs) while Swift just uses HTTP. Also performance considerations are not discussed here.
As a comparison, the Linux kernel has roughly 25 million lines of code and a tool like GNU make has about 33000 lines of code. Make is not a very complex piece of software. Is OpenStack swift?





Saturday, May 12, 2012

Starting to research OpenStack Swift

As we are always looking at lowering our storage costs while still trying to manage petabytes of storage we heard about "object storage" for a few years. This Buzzword sounds a bit like a bad disease to a traditional Linux/Unix heavy Scientific Computing shop. It sounds like something that could break in all sorts of ways and would have unbearable latency etc.


On the other hand we see almost every day that storage and other IT vendors are jumping on the object and cloud storage bandwagon. Is it all just cloud hype or is there something more to it? One platform that sticks out particularly is OpenStack after more than a dozen companies (AT&T, IBM, 
Red Hat, SUSE, Cisco, Dell, Canonical, etc) have pledged to support the OpenStack foundation. OpenStack was created by Rackspace and NASA (here is the story behind it) and the storage component Swift was originally developed at Rackspace. As we are most interested in storage, Swift is the thing we are looking at. 
Now, is this really a OSS project with broad support and many contributors? Until today Rackspace appears to be doing most of the real work, but there is a fair number of other big names who are also contributing code.


We work quite a bit with Dell hardware and it is nice to see that they have created a nice deployment solution called Crowbar that uses an OSS DevOps approach to push openstack to their servers. Their cloud dude seems to be a bit of an OpenStack enthusiast. But there are also a few startups that are betting on OpenStack Swift, such as SwiftStack.com who sells you a customized Ubuntu Image with a web management tool that lets you deploy a Swift storage cluster in a few minutes. The SwiftStack people are core contributors to the OpenStack swift project so they know the code base very well.
How about end user adoption in Universities and other research places? The San Diego Super Computing Center has brought their OpenStack storage cloud online last year and is offering pretty reasonable pricing (about 1/3 of the price of S3).
Why are all these large companies joining OpenStack? Well, of course they all are way behind Amazon EC2/S3 and joining forces can either be seen as a good strategy or as a desperate attempt to catch up. 

From a storage technology perspective there are may be 3 reasons for this push that come to my mind. First, it takes a very long time to develop a storage platform. For BlueArc, 3PAR, Compellent, Isilon, etc it took almost 10 years to convince many IT managers that those were viable options. HP and Dell needed to suck up one of those manufacturers to get the know how. Second, customers are increasing vary of vendor lock in and lack of scalability because big data capacity and especially performance needs are very  unpredictable. And third, traditional storage techniques such as RAID will not be viable
in the future and alternatives (examples are gpfs, panassas but also 3PAR with it's chunklet stuff) take a very long time to develop (again, see first point).

But why does OpenStack seem to have more followers than CloudStack, Eucalyptus or others? It is extremely scalable but I could not (yet) find any strong hints that it is more scalable than other stacks.
From a developer and system integrator view the OpenStack trump card seems to be modularity which is important for keeping up development speed and for allowing a large community of developers to participate. 

What strikes me from a systems management perspective is the simplicity of the underlying toolset. Every Unix admin is familiar with Python, Sqlite, Rsync and Linux/XFS. At first you might think: What, that's what they are using? After all, rsync is more than 15 years old and this is the tool that is supposed to help conquering the storage world in the 21st century?
Then you think: Oh if our sysadmins ever have to do a root cause analysis on performance issues they already know rsync and if they ever have to throttle the replication engine they already know what --bwlimit is. That does not sound too bad....but we will have to take a deeper look at this ..... to be continued.




Random Links & Blogs:
http://programmerthoughts.com/openstack/swift-tech-overview/
http://searchstorage.techtarget.com/news/2240105808/Caringo-CAStor-integrates-object-storage-with-OpenStack-Swift
http://www.slideshare.net/HuiCheng2/integrating-open-stack
http://www.buildcloudstorage.com/
http://www.cloudconnectevent.com/santaclara/2012/presentations/free/99-john-dickinson.pdf
http://www.buildcloudstorage.com/2012/01/can-openstack-swift-hit-amazon-s3-like.html

Consultants:
http://www.talkincloud.com/it-consultants-build-openstack-cloud-business-practices/
http://www.griddynamics.com/ or http://openstackgd.wordpress.com/