Friday, February 27, 2009

How to resize ZFS - Part 2 (the real world)

As you can read here How to resize ZFS, it is possible to resize a zpool if you replace each disk of the pool with a bigger one. In my first post I used a virtual machine to test this. And... I've replaced one disk (eg. da0) with a different device (eg. da4).
In the real world I would like to replace one device with the same (but bigger one). I was able to get some very old SCSI Disks in an external SCSI enclosure. These disks are perfect for testing this...

Here we go...

My testsystem is a Sempron 2600+ with 1 GB RAM and several SCSI Disks connected via a DAWICONTROL DC-2976 UW controller. The FreeNAS Version I've used for this test is 0.7 Sardaukar (revision 4390)

I've created a testpool with three 2.1 GB SCSI-Disks (yes, it's true, just 2.1 GB ;-) ). And I've also created two filesystems (testfs1, compressed & testfs2), and some testfiles.

crusher:~# zpool list testpool0
NAME SIZE USED AVAIL CAP HEALTH ALTROOT
testpool0 5.91G 2.81G 3.09G 47% ONLINE -

As you can see, this pool has a size of 5.91 GB (remember this)

crusher:~# zfs list
NAME USED AVAIL REFER MOUNTPOINT
bigpool 100M 364G 25.3K /mnt/bigpool
bigpool/temp 100M 364G 100M /mnt/bigpool/temp
testpool0 1.87G 2.00G 28.0K /mnt/testpool0
testpool0/testfs1 640M 2.00G 640M /mnt/testpool0/testfs1
testpool0/testfs2 1.25G 2.00G 1.25G /mnt/testpool0/testfs2

crusher:~# zpool status testpool0
pool: testpool0
state: ONLINE
scrub: none requested
config:

NAME STATE READ WRITE CKSUM
testpool0 ONLINE 0 0 0
raidz1 ONLINE 0 0 0
da0 ONLINE 0 0 0
da1 ONLINE 0 0 0
da2 ONLINE 0 0 0

errors: No known data errors

First, I'd like to start with a scrub to be sure that the data is safe on these disks...

crusher:~# zpool scrub testpool

This will take a while. Check the status with:

crusher:~# zpool status testpool0
pool: testpool0
state: ONLINE
scrub: scrub in progress, 93.34% done, 0h0m to go
config:

NAME STATE READ WRITE CKSUM
testpool0 ONLINE 0 0 0
raidz1 ONLINE 0 0 0
da0 ONLINE 0 0 0
da1 ONLINE 0 0 0
da2 ONLINE 0 0 0

errors: No known data errors

crusher:~# zpool status testpool0
pool: testpool0
state: ONLINE
scrub: scrub completed with 0 errors on Fri Feb 27 21:40:16 2009
config:

NAME STATE READ WRITE CKSUM
testpool0 ONLINE 0 0 0
raidz1 ONLINE 0 0 0
da0 ONLINE 0 0 0
da1 ONLINE 0 0 0
da2 ONLINE 0 0 0

errors: No known data errors

Everything looks perfect... so lets start with the replace. This SCSI enclosure supports Hot-Swapping of the disks (more or less).

Ok, lets replace the first disk (da0)

Offline the first disk

crusher:~# zpool offline testpool0 da0
Bringing device da0 offline
crusher:~# zpool status testpool0
pool: testpool0
state: DEGRADED
status: One or more devices has been taken offline by the administrator.
Sufficient replicas exist for the pool to continue functioning in a
degraded state.
action: Online the device using 'zpool online' or replace the device with
'zpool replace'.
scrub: scrub completed with 0 errors on Fri Feb 27 21:40:16 2009
config:

NAME STATE READ WRITE CKSUM
testpool0 DEGRADED 0 0 0
raidz1 DEGRADED 0 0 0
da0 OFFLINE 0 0 0
da1 ONLINE 0 0 0
da2 ONLINE 0 0 0

errors: No known data errors

Replace the first disk (da0)

First physicaly ;-)

And then do the replace command

crusher:~# zpool replace testpool0 da0
crusher:~# zpool status testpool0
pool: testpool0
state: DEGRADED
status: One or more devices is currently being resilvered. The pool will
continue to function, possibly in a degraded state.
action: Wait for the resilver to complete.
scrub: resilver in progress, 14.79% done, 0h2m to go
config:

NAME STATE READ WRITE CKSUM
testpool0 DEGRADED 0 0 0
raidz1 DEGRADED 0 0 0
replacing DEGRADED 0 0 0
da0/old OFFLINE 0 0 0
da0 ONLINE 0 0 0
da1 ONLINE 0 0 0
da2 ONLINE 0 0 0

errors: No known data errors

After a while, the resivering has finished

crusher:~# zpool status testpool
pool: testpool
state: ONLINE
scrub: resilver completed with 0 errors on Fri Feb 27 20:05:52 2009
config:

NAME STATE READ WRITE CKSUM
testpool ONLINE 0 0 0
raidz1 ONLINE 0 0 0
da0 ONLINE 0 0 0
da1 ONLINE 0 0 0
da2 ONLINE 0 0 0

errors: No known data errors

Lets replace the second and third disk...

crusher:~# zpool status testpool0
pool: testpool0
state: ONLINE
scrub: scrub completed with 0 errors on Fri Feb 27 22:57:19 2009
config:

NAME STATE READ WRITE CKSUM
testpool0 ONLINE 0 0 0
raidz1 ONLINE 0 0 0
da0 ONLINE 0 0 0
da1 ONLINE 0 0 0
da2 ONLINE 0 0 0

errors: No known data errors

crusher:~# zpool list
NAME SIZE USED AVAIL CAP HEALTH ALTROOT
bigpool 556G 150M 556G 0% ONLINE -
testpool0 5.91G 2.82G 3.09G 47% ONLINE -

crusher:~# zfs list
NAME USED AVAIL REFER MOUNTPOINT
bigpool 100M 364G 25.3K /mnt/bigpool
bigpool/temp 100M 364G 100M /mnt/bigpool/temp
testpool0 1.87G 2.00G 28.0K /mnt/testpool0
testpool0/testfs1 640M 2.00G 640M /mnt/testpool0/testfs1
testpool0/testfs2 1.25G 2.00G 1.25G /mnt/testpool0/testfs2

As you can see the size of the testpool hasn't changed. After exporting and importing the pool, the additional space is available!

crusher:~# zpool export testpool0
crusher:~# zpool import testpool0
crusher:~# zpool list
NAME SIZE USED AVAIL CAP HEALTH ALTROOT
bigpool 556G 150M 556G 0% ONLINE -
testpool0 12.0G 2.82G 9.15G 23% ONLINE -

The size has been changed to 12.0 GB...

Tuesday, January 13, 2009

FreeNAS 0.7 (rev. 3953) and ZFS - Update 2

In my previous post, I've written that the workaround for FreeNAS 0.7 rev. 3953 works, but after a reboot, the zpool is gone and has to be imported by hand (zpool import -f ...)
Stefan left a comment that his full installation works perfect...

I can confirm this! With a full installation and the workaround, the pool is available after a reboot... It's getting better :-)


Monday, January 12, 2009

Blank smbpasswd - FreeNAS


While I'm hitting this issue (again :-) ), so I want to document this here...

My problem was that I had a blank smbpasswd. Even when I added the passwd manually with 'smbpasswd'. After a reboot, the password settings in /var/etc/private/smbpasswd are gone.
With debugging /etc/rc.d/smbpasswd (with set -x) I've found a variable called 'smbpasswd_minuid=1001'. My UID's are bellow 1001 (I'm using the same as on my Mac).
So after some search, I've seen that this feature (Enable customizing of minimum UID for smbpasswd via rc.conf variable 'smbpasswd_minuid'.) was implemented with Build 3177.

It is easy to set this variable in -> System -> Advanced -> rc.conf

Variable = smbpasswd_minuid  ; Value = 500


After applying this, it is possible to run /etc/rc.d/smbpasswd. The users and the encrypted passwords are stored in /var/etc/private/smbpasswd.

Sunday, January 11, 2009

FreeNAS 0.7 (rev. 3953) and ZFS - Update

As I've written here, ZFS is not working with the latest available build (i386, 3953) of FreeNAS 0.7.

There is a workaround described in the forum (Re: ZFS issues? kernel: KLD zfs.ko: depends on opensolaris-not a)
I've tested this workaround, and it works... until the reboot :-(

After the reboot the zpool is gone and you have to import it again, or export it 'before' you want to reboot your NAS.

I'm really looking forward for a new build...

Thursday, December 18, 2008

Backup FreeNAS using rsync to a rsync-server

It is a little bit difficult to find out how to backup data from a FreeNAS-Server to a different rsync-server. In my case I have a QNAP TS-209 configured as a rsync-server. To this NAS device, I want to backup my data once a day.

The trick is to use the LOCAL rsync configuration of FreeNAS. The rsync-server is here called 'soulcube'. 


I've configured four backup jobs. Here is the example of my photo share.


Source share - the_directory_you_want_to_backup
Destination share - servername::rsync_share

Tuesday, December 9, 2008

FreeNAS 0.7 (rev. 3953) and ZFS

I've tried to use with the latest Nightly Buid of FreeNAS 0.7 (rev. 3953) ZFS. Unfortunately I was not able to import my ZFS-Pool.

The error message just showed the following...

kernel: KLD zfs.ko: depends on opensolaris - not available

Currently there is no solution for this problem :-(

Tuesday, October 14, 2008

Benchmarking FreeNAS WritePerformance

Jonas left me a comment here about strange write performance of his FreeNAS 0.7. 
Jonas said...

Im haveing issues with performance on raidz for Freenas. Right now i have four 500GB disks in a raidz1 pool.

Trying to dd 1GB file to my pool is only giving a write performance of 47MB/s. If i test one single disk with diskinfo -tv i get write 48-82MB.

After some private mail and some testing, we've found out that this has something to do with the dd options Jonas used.

freenas:/mnt/Media# dd if=/dev/zero of=mytestfile.out bs=1000 count=1000000 
1000000+0 records in 
1000000+0 records out 
1000000000 bytes transferred in 21.231106 secs (47100702 bytes/sec)


Around 47 MB/s for a very performant system is not enough (AMD Athlon x2 64 Bit 4200+, 2200 MHz, 2 GByte RAM, 4x 500 GB Harddisks configured as RAIDZ1)

Speed of a single disk was measured with diskinfo

freenas:/mnt/Media# diskinfo -tv ad4 
ad4 
512 # sectorsize 
500107862016 # mediasize in bytes (466G) 
976773168 # mediasize in sectors 
969021 # Cylinders according to firmware. 
16 # Heads according to firmware. 
63 # Sectors according to firmware. 
ad:S13TJ1MQ702083 # Disk ident. 
 
Seek times: 
Full stroke: 250 iter in 5.368466 sec = 21.474 msec 
Half stroke: 250 iter in 3.997566 sec = 15.990 msec 
Quarter stroke: 500 iter in 6.592644 sec = 13.185 msec 
Short forward: 400 iter in 2.437786 sec = 6.094 msec 
Short backward: 400 iter in 1.262446 sec = 3.156 msec 
Seq outer: 2048 iter in 0.251187 sec = 0.123 msec 
Seq inner: 2048 iter in 0.243935 sec = 0.119 msec 
Transfer rates: 
outside: 102400 kbytes in 1.245477 sec = 82217 kbytes/sec 
middle: 102400 kbytes in 1.417240 sec = 72253 kbytes/sec 
inside: 102400 kbytes in 2.351795 sec = 43541 kbytes/sec


This looks OK. So whats the problem?

It is the blocksize (bs) dd used to write the testfile! It's just 1000 Bytes!

Use higher numbers and a 'binary prefix' (see http://en.wikipedia.org/wiki/Binary_prefix). I mean something like 512k, 1024k, 2048k, 4096k, 8192k.

Jonas repeated the test with the following results:

freenas:/mnt/Media# dd if=/dev/zero of=testfile bs=8192k count=100
100+0 records in
100+0 records out
838860800 bytes transferred in 5.715633 secs (146766032 bytes/sec)

freenas:/mnt/Media# dd if=/dev/zero of=testfile bs=4096k count=200
200+0 records in
200+0 records out
838860800 bytes transferred in 5.663804 secs (148109079 bytes/sec)

freenas:/mnt/Media# dd if=/dev/zero of=testfile bs=2048k count=200
200+0 records in
200+0 records out
419430400 bytes transferred in 3.429654 secs (122295248 bytes/sec)

freenas:/mnt/Media# dd if=/dev/zero of=testfile bs=1024k count=200
200+0 records in
200+0 records out
209715200 bytes transferred in 0.357848 secs (586045588 bytes/sec)  << To little data gives strange result. Probably because of HD cache.

freenas:/mnt/Media# dd if=/dev/zero of=testfile bs=512k count=1000
1000+0 records in
1000+0 records out
524288000 bytes transferred in 2.576585 secs (203481736 bytes/sec)

freenas:/mnt/Media# dd if=/dev/zero of=testfile bs=256k count=10000
10000+0 records in
10000+0 records out
2621440000 bytes transferred in 16.343732 secs (160394212 bytes/sec)

freenas:/mnt/Media# dd if=/dev/zero of=testfile bs=25k count=100000
100000+0 records in
100000+0 records out
2560000000 bytes transferred in 20.552446 secs (124559383 bytes/sec)


So, it depends how you measure the performance ;-)