--- /dev/null
+Return-Path: <bremner@unb.ca>\r
+X-Original-To: notmuch@notmuchmail.org\r
+Delivered-To: notmuch@notmuchmail.org\r
+Received: from localhost (localhost [127.0.0.1])\r
+ by olra.theworths.org (Postfix) with ESMTP id 46F99431FB6\r
+ for <notmuch@notmuchmail.org>; Wed, 20 Feb 2013 17:29:44 -0800 (PST)\r
+X-Virus-Scanned: Debian amavisd-new at olra.theworths.org\r
+X-Spam-Flag: NO\r
+X-Spam-Score: 0\r
+X-Spam-Level: \r
+X-Spam-Status: No, score=0 tagged_above=-999 required=5 tests=[none]\r
+ autolearn=disabled\r
+Received: from olra.theworths.org ([127.0.0.1])\r
+ by localhost (olra.theworths.org [127.0.0.1]) (amavisd-new, port 10024)\r
+ with ESMTP id eid1EeS5z6Da for <notmuch@notmuchmail.org>;\r
+ Wed, 20 Feb 2013 17:29:43 -0800 (PST)\r
+Received: from tesseract.cs.unb.ca (tesseract.cs.unb.ca [131.202.240.238])\r
+ (using TLSv1 with cipher DHE-RSA-AES128-SHA (128/128 bits))\r
+ (No client certificate requested)\r
+ by olra.theworths.org (Postfix) with ESMTPS id 346DE431FAE\r
+ for <notmuch@notmuchmail.org>; Wed, 20 Feb 2013 17:29:43 -0800 (PST)\r
+Received: from fctnnbsc30w-156034082078.dhcp-dynamic.fibreop.nb.bellaliant.net\r
+ ([156.34.82.78] helo=zancas.localnet)\r
+ by tesseract.cs.unb.ca with esmtpsa\r
+ (TLS1.2:DHE_RSA_AES_128_CBC_SHA1:128) (Exim 4.80)\r
+ (envelope-from <bremner@unb.ca>)\r
+ id 1U8Kyg-0000v5-74; Wed, 20 Feb 2013 21:29:38 -0400\r
+Received: from bremner by zancas.localnet with local (Exim 4.80)\r
+ (envelope-from <bremner@unb.ca>)\r
+ id 1U8KyZ-0007UO-1d; Wed, 20 Feb 2013 21:29:31 -0400\r
+From: David Bremner <david@tethera.net>\r
+To: notmuch mailing list <notmuch@notmuchmail.org>\r
+Subject: Re: On disk tag storage format\r
+In-Reply-To: <874nk8v9zw.fsf@zancas.localnet>\r
+References: <874nk8v9zw.fsf@zancas.localnet>\r
+User-Agent: Notmuch/0.15.2+32~g16aa65b (http://notmuchmail.org) Emacs/24.2.1\r
+ (x86_64-pc-linux-gnu)\r
+Date: Wed, 20 Feb 2013 21:29:30 -0400\r
+Message-ID: <87vc9mtpxh.fsf@zancas.localnet>\r
+MIME-Version: 1.0\r
+Content-Type: multipart/mixed; boundary="=-=-="\r
+X-Spam_bar: -\r
+X-BeenThere: notmuch@notmuchmail.org\r
+X-Mailman-Version: 2.1.13\r
+Precedence: list\r
+List-Id: "Use and development of the notmuch mail system."\r
+ <notmuch.notmuchmail.org>\r
+List-Unsubscribe: <http://notmuchmail.org/mailman/options/notmuch>,\r
+ <mailto:notmuch-request@notmuchmail.org?subject=unsubscribe>\r
+List-Archive: <http://notmuchmail.org/pipermail/notmuch>\r
+List-Post: <mailto:notmuch@notmuchmail.org>\r
+List-Help: <mailto:notmuch-request@notmuchmail.org?subject=help>\r
+List-Subscribe: <http://notmuchmail.org/mailman/listinfo/notmuch>,\r
+ <mailto:notmuch-request@notmuchmail.org?subject=subscribe>\r
+X-List-Received-Date: Thu, 21 Feb 2013 01:29:44 -0000\r
+\r
+--=-=-=\r
+Content-Type: text/plain\r
+\r
+David Bremner <david@tethera.net> writes:\r
+\r
+> Austin outlined on IRC a way of representing tags on disk as hardlinks\r
+> to messages. In order to make the discussion more concrete, I wrote a\r
+> prototype in python to dump the notmuch database to this format. On my\r
+> 250k messages, this creates 40k new hardlinks, and uses about 5M of\r
+> diskspace. The dump process takes about 20s on\r
+> my core i7 machine. With symbolic links, the same database takes about\r
+> 150M of disk space; this isn't great but it isn't unbearable either.\r
+>\r
+\r
+I've being playing a bit with this script and it seems more or less\r
+usable as a way of mirroring the notmuch tag database to a link farm.\r
+\r
+It's a bit faster than my current dump/restore based approach, although\r
+if you want to keep the results in a git repository then it takes up\r
+more space. Of course the bonus with this approach is that it creates\r
+"virtual" maildirs for each tag that can be browsed with the maildir\r
+client of choice.\r
+\r
+The current default is to use some mix of hard and symbolic links to try\r
+to balance the space consumed in a git repo versus the inode\r
+consumption/performance issues of using too many symlinks.\r
+\r
+It's still a prototype, and there is not much error checking, and there\r
+are certain issues not dealt with at all (the ones I thought about are\r
+commented).\r
+\r
+\r
+--=-=-=\r
+Content-Type: text/x-python\r
+Content-Disposition: inline; filename=linksync.py\r
+\r
+# Copyright 2013, David Bremner <david@tethera.net>\r
+\r
+# Licensed under the same terms as notmuch.\r
+\r
+import notmuch\r
+import re\r
+import os, errno\r
+import sys\r
+from collections import defaultdict\r
+import argparse\r
+\r
+# skip automatic and maildir tags\r
+\r
+skiptags = re.compile(r"^(attachement|signed|encrypted|draft|flagged|passed|replied|unread)$")\r
+\r
+# some random person on stack overflow suggests:\r
+\r
+def mkdir_p(path):\r
+ try:\r
+ os.makedirs(path)\r
+ except OSError as exc: # Python >2.5\r
+ if exc.errno == errno.EEXIST and os.path.isdir(path):\r
+ pass\r
+ else: raise\r
+\r
+CHARSET = 'ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+_@=.,-'\r
+\r
+encode_re = '([^{0}])'.format(CHARSET)\r
+\r
+decode_re = '[%]([0-7][0-9A-Fa-f])'\r
+\r
+def encode_one_char(match):\r
+ return('%{:02x}'.format(ord(match.group(1))))\r
+\r
+def encode_for_fs(str):\r
+ return re.sub(encode_re,encode_one_char, str,0)\r
+\r
+def decode_one_char(match):\r
+ return chr(int(match.group(1),16))\r
+\r
+def decode_from_fs(str):\r
+ return re.sub(decode_re,decode_one_char, str, 0)\r
+\r
+\r
+def mk_tag_dir(tagdir):\r
+\r
+ mkdir_p (os.path.join(tagdir, 'cur'))\r
+ mkdir_p (os.path.join(tagdir, 'new'))\r
+ mkdir_p (os.path.join(tagdir, 'tmp'))\r
+\r
+\r
+flagpart = '(:2,[^:]*)'\r
+flagre = re.compile(flagpart + '$');\r
+\r
+def path_for_msg (dir, msg):\r
+ filename = msg.get_filename()\r
+ flagsmatch = flagre.search(filename)\r
+ if flagsmatch == None:\r
+ flags = ''\r
+ else:\r
+ flags = flagsmatch.group(1)\r
+\r
+ return os.path.join(dir, 'cur', encode_for_fs(msg.get_message_id()) + flags)\r
+\r
+\r
+def unlink_message(dir, msg):\r
+\r
+ dir = os.path.join(dir, 'cur')\r
+\r
+ filepattern = encode_for_fs(msg.get_message_id()) + flagpart +'?$'\r
+\r
+ filere = re.compile(filepattern);\r
+\r
+ for file in os.listdir(dir):\r
+ if filere.match(file):\r
+ os.unlink(os.path.join(dir, file))\r
+\r
+def dir_for_tag(tag):\r
+ enc_tag = encode_for_fs (tag)\r
+ return os.path.join(tagroot, enc_tag)\r
+\r
+disk_tags = defaultdict(set)\r
+disk_ids = set()\r
+\r
+def read_tags_from_disk(rootdir):\r
+\r
+ for root, subFolders, files in os.walk(rootdir):\r
+ for filename in files:\r
+ msg_id = filename.split(':')[0]\r
+ tag = root.split('/')[-2]\r
+ decoded_id = decode_from_fs(msg_id)\r
+ disk_ids.add(decoded_id)\r
+ disk_tags[decoded_id].add(decode_from_fs(tag));\r
+\r
+# Main program\r
+\r
+parser = argparse.ArgumentParser(description='Sync notmuch tag database to/from link farm')\r
+parser.add_argument('-l','--link-style',choices=['hard','symbolic', 'adaptive'],\r
+ default='adaptive',dest='link_style')\r
+parser.add_argument('-d','--destination',choices=['disk','notmuch'], default='disk',\r
+ dest='destination')\r
+parser.add_argument('-t','--threshold', default=50000L, type=int, dest='threshold')\r
+\r
+parser.add_argument('tagroot')\r
+\r
+opts=parser.parse_args()\r
+\r
+tagroot=opts.tagroot\r
+\r
+sync_from_links = (opts.destination == 'notmuch')\r
+\r
+read_tags_from_disk(tagroot)\r
+\r
+if sync_from_links:\r
+ db = notmuch.Database(mode=notmuch.Database.MODE.READ_WRITE)\r
+else:\r
+ db = notmuch.Database(mode=notmuch.Database.MODE.READ_ONLY)\r
+\r
+dbtags = filter (lambda tag: not skiptags.match(tag), db.get_all_tags())\r
+\r
+querystr = ' OR '.join(map (lambda tag: 'tag:'+tag, dbtags));\r
+\r
+q_new = notmuch.Query(db, querystr)\r
+q_new.set_sort(notmuch.Query.SORT.UNSORTED)\r
+for msg in q_new.search_messages():\r
+\r
+ # silently ignore empty tags\r
+ db_tags = set(filter (lambda tag: tag != '' and not skiptags.match(tag),\r
+ msg.get_tags()))\r
+\r
+ message_id = msg.get_message_id()\r
+\r
+ disk_ids.discard(message_id)\r
+\r
+ missing_on_disk = db_tags.difference(disk_tags[message_id])\r
+ missing_in_db = disk_tags[message_id].difference(db_tags)\r
+\r
+ if sync_from_links:\r
+ msg.freeze()\r
+\r
+ filename = msg.get_filename()\r
+\r
+ if len(missing_on_disk) > 0:\r
+ if opts.link_style == 'adaptive':\r
+ statinfo = os.stat (filename)\r
+ symlink = (statinfo.st_size > opts.threshold)\r
+ else:\r
+ symlink = opts.link_style == 'symbolic'\r
+\r
+ for tag in missing_on_disk:\r
+\r
+ if sync_from_links:\r
+ msg.remove_tag(tag,sync_maildir_flags=False)\r
+ else:\r
+ tagdir = dir_for_tag (tag)\r
+ mk_tag_dir (tagdir)\r
+\r
+ newlink = path_for_msg (tagdir, msg)\r
+\r
+ if symlink:\r
+ os.symlink(filename, newlink)\r
+ else:\r
+ os.link(filename, newlink)\r
+\r
+\r
+ for tag in missing_in_db:\r
+ if sync_from_links:\r
+ msg.add_tag(tag,sync_maildir_flags=False)\r
+ else:\r
+ tagdir = dir_for_tag (tag)\r
+ unlink_message(tagdir,msg)\r
+\r
+ if sync_from_links:\r
+ msg.thaw()\r
+\r
+# everything remaining in disk_ids is a deleted message\r
+# unless we are syncing back to the database, in which case\r
+# it just might not currently have any non maildir tags.\r
+\r
+if not sync_from_links:\r
+ for root, subFolders, files in os.walk(tagroot):\r
+ for filename in files:\r
+ msg_id = filename.split(':')[0]\r
+ decoded_id = decode_from_fs(msg_id)\r
+ if decoded_id in disk_ids:\r
+ os.unlink(os.path.join(root, filename))\r
+\r
+\r
+db.close()\r
+\r
+# currently empty directories are not pruned.\r
+\r
+--=-=-=--\r