Re: [Patch v2 01/17] hex-escape: (en|de)code strings to/from restricted character set
authorJani Nikula <jani@nikula.org>
Fri, 30 Nov 2012 21:43:38 +0000 (23:43 +0200)
committerW. Trevor King <wking@tremily.us>
Fri, 7 Nov 2014 17:51:14 +0000 (09:51 -0800)
ac/b9f34c163752d2ef49b8c4df483d7828ef5eb9 [new file with mode: 0644]

diff --git a/ac/b9f34c163752d2ef49b8c4df483d7828ef5eb9 b/ac/b9f34c163752d2ef49b8c4df483d7828ef5eb9
new file mode 100644 (file)
index 0000000..54c4756
--- /dev/null
@@ -0,0 +1,351 @@
+Return-Path: <jani@nikula.org>\r
+X-Original-To: notmuch@notmuchmail.org\r
+Delivered-To: notmuch@notmuchmail.org\r
+Received: from localhost (localhost [127.0.0.1])\r
+       by olra.theworths.org (Postfix) with ESMTP id 947D4431FAF\r
+       for <notmuch@notmuchmail.org>; Fri, 30 Nov 2012 13:43:49 -0800 (PST)\r
+X-Virus-Scanned: Debian amavisd-new at olra.theworths.org\r
+X-Spam-Flag: NO\r
+X-Spam-Score: -0.7\r
+X-Spam-Level: \r
+X-Spam-Status: No, score=-0.7 tagged_above=-999 required=5\r
+       tests=[RCVD_IN_DNSWL_LOW=-0.7] autolearn=disabled\r
+Received: from olra.theworths.org ([127.0.0.1])\r
+       by localhost (olra.theworths.org [127.0.0.1]) (amavisd-new, port 10024)\r
+       with ESMTP id iEVbKlHTBOyl for <notmuch@notmuchmail.org>;\r
+       Fri, 30 Nov 2012 13:43:45 -0800 (PST)\r
+Received: from mail-la0-f53.google.com (mail-la0-f53.google.com\r
+       [209.85.215.53]) (using TLSv1 with cipher RC4-SHA (128/128 bits))\r
+       (No client certificate requested)\r
+       by olra.theworths.org (Postfix) with ESMTPS id 67A9E431FAE\r
+       for <notmuch@notmuchmail.org>; Fri, 30 Nov 2012 13:43:44 -0800 (PST)\r
+Received: by mail-la0-f53.google.com with SMTP id w12so789442lag.26\r
+       for <notmuch@notmuchmail.org>; Fri, 30 Nov 2012 13:43:42 -0800 (PST)\r
+X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed;\r
+       d=google.com; s=20120113;\r
+       h=from:to:cc:subject:in-reply-to:references:user-agent:date\r
+       :message-id:mime-version:content-type:x-gm-message-state;\r
+       bh=vaJl32SvhOshPg0lRn8ibEGhmC3r1fvHvNc6sEiaCw8=;\r
+       b=FNSd95Je10y+cF6gmh3JPWNq1vbxs6OD61GJImkBFJjm+BatNTuaX5HKIk9ZEJyfYf\r
+       PCILp35n9w1Cn0ojED7Jwet4TyZZP7AkZpH7hZZd5K9vDFnqCMnmajfMPSJmE2QWiVF/\r
+       bsSY1AbSTIR2HqrGM11eDe7aoHOGgbvGCrFtfmoeZ/cNrhFzbHGwDU+7/rLCsaBOswl2\r
+       7UtG1V0nevlUQxaxjFQB4q33HeLGQjqtJoHVD6KASHtAFUlAdgAD3sg93anhs3/gKfwK\r
+       PQMoMSZ3VH0QEE8bXN02IWHTdH8ftXA25TlgqFFkdSVYJ5pGD09jUiWJ9PFrFwn91SFU\r
+       dhKg==\r
+Received: by 10.112.83.133 with SMTP id q5mr1424247lby.40.1354311822497;\r
+       Fri, 30 Nov 2012 13:43:42 -0800 (PST)\r
+Received: from localhost (dsl-hkibrasgw4-fe51df00-27.dhcp.inet.fi.\r
+       [80.223.81.27])\r
+       by mx.google.com with ESMTPS id y10sm2519789lbg.4.2012.11.30.13.43.40\r
+       (version=SSLv3 cipher=OTHER); Fri, 30 Nov 2012 13:43:41 -0800 (PST)\r
+From: Jani Nikula <jani@nikula.org>\r
+To: david@tethera.net, notmuch@notmuchmail.org\r
+Subject: Re: [Patch v2 01/17] hex-escape: (en|de)code strings to/from\r
+       restricted character set\r
+In-Reply-To: <1353792017-31459-2-git-send-email-david@tethera.net>\r
+References: <1353792017-31459-1-git-send-email-david@tethera.net>\r
+       <1353792017-31459-2-git-send-email-david@tethera.net>\r
+User-Agent: Notmuch/0.14+124~g3b17402 (http://notmuchmail.org) Emacs/23.4.1\r
+       (i686-pc-linux-gnu)\r
+Date: Fri, 30 Nov 2012 23:43:38 +0200\r
+Message-ID: <87wqx2ix6d.fsf@nikula.org>\r
+MIME-Version: 1.0\r
+Content-Type: text/plain; charset=us-ascii\r
+X-Gm-Message-State:\r
+ ALoCoQkPIg2uqputwiYaIbCItPz1RTVWy81b65SwMZ7wTonNRXDIfJbuu2+rhWn/SnmdUzuYs2gg\r
+Cc: David Bremner <bremner@debian.org>\r
+X-BeenThere: notmuch@notmuchmail.org\r
+X-Mailman-Version: 2.1.13\r
+Precedence: list\r
+List-Id: "Use and development of the notmuch mail system."\r
+       <notmuch.notmuchmail.org>\r
+List-Unsubscribe: <http://notmuchmail.org/mailman/options/notmuch>,\r
+       <mailto:notmuch-request@notmuchmail.org?subject=unsubscribe>\r
+List-Archive: <http://notmuchmail.org/pipermail/notmuch>\r
+List-Post: <mailto:notmuch@notmuchmail.org>\r
+List-Help: <mailto:notmuch-request@notmuchmail.org?subject=help>\r
+List-Subscribe: <http://notmuchmail.org/mailman/listinfo/notmuch>,\r
+       <mailto:notmuch-request@notmuchmail.org?subject=subscribe>\r
+X-List-Received-Date: Fri, 30 Nov 2012 21:43:49 -0000\r
+\r
+On Sat, 24 Nov 2012, david@tethera.net wrote:\r
+> From: David Bremner <bremner@debian.org>\r
+>\r
+> The character set is chosen to be suitable for pathnames, and the same\r
+> as that used by contrib/nmbug\r
+>\r
+> [With additions by Jani Nikula]\r
+\r
+So it must be good. ;)\r
+\r
+Just a couple of nitpicks below.\r
+\r
+BR,\r
+Jani.\r
+\r
+> ---\r
+>  util/Makefile.local |    2 +-\r
+>  util/hex-escape.c   |  168 +++++++++++++++++++++++++++++++++++++++++++++++++++\r
+>  util/hex-escape.h   |   41 +++++++++++++\r
+>  3 files changed, 210 insertions(+), 1 deletion(-)\r
+>  create mode 100644 util/hex-escape.c\r
+>  create mode 100644 util/hex-escape.h\r
+>\r
+> diff --git a/util/Makefile.local b/util/Makefile.local\r
+> index c7cae61..3ca623e 100644\r
+> --- a/util/Makefile.local\r
+> +++ b/util/Makefile.local\r
+> @@ -3,7 +3,7 @@\r
+>  dir := util\r
+>  extra_cflags += -I$(srcdir)/$(dir)\r
+>  \r
+> -libutil_c_srcs := $(dir)/xutil.c $(dir)/error_util.c\r
+> +libutil_c_srcs := $(dir)/xutil.c $(dir)/error_util.c $(dir)/hex-escape.c\r
+>  \r
+>  libutil_modules := $(libutil_c_srcs:.c=.o)\r
+>  \r
+> diff --git a/util/hex-escape.c b/util/hex-escape.c\r
+> new file mode 100644\r
+> index 0000000..d8905d0\r
+> --- /dev/null\r
+> +++ b/util/hex-escape.c\r
+> @@ -0,0 +1,168 @@\r
+> +/* hex-escape.c -  Manage encoding and decoding of byte strings into path names\r
+> + *\r
+> + * Copyright (c) 2011 David Bremner\r
+> + *\r
+> + * This program is free software: you can redistribute it and/or modify\r
+> + * it under the terms of the GNU General Public License as published by\r
+> + * the Free Software Foundation, either version 3 of the License, or\r
+> + * (at your option) any later version.\r
+> + *\r
+> + * This program is distributed in the hope that it will be useful,\r
+> + * but WITHOUT ANY WARRANTY; without even the implied warranty of\r
+> + * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE.  See the\r
+> + * GNU General Public License for more details.\r
+> + *\r
+> + * You should have received a copy of the GNU General Public License\r
+> + * along with this program.  If not, see http://www.gnu.org/licenses/ .\r
+> + *\r
+> + * Author: David Bremner <david@tethera.net>\r
+> + */\r
+> +\r
+> +#include <assert.h>\r
+> +#include <string.h>\r
+> +#include <talloc.h>\r
+> +#include <ctype.h>\r
+> +#include "error_util.h"\r
+> +#include "hex-escape.h"\r
+> +\r
+> +static const size_t default_buf_size = 1024;\r
+> +\r
+> +static const char *output_charset =\r
+> +    "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz0123456789+-_@=.:,";\r
+> +\r
+> +static const char escape_char = '%';\r
+> +\r
+> +static int\r
+> +is_output (char c)\r
+> +{\r
+> +    return (strchr (output_charset, c) != NULL);\r
+> +}\r
+> +\r
+> +static int\r
+> +maybe_realloc (void *ctx, size_t needed, char **out, size_t *out_size)\r
+> +{\r
+> +    if (*out_size < needed) {\r
+> +\r
+> +    if (*out == NULL)\r
+> +        *out = talloc_size (ctx, needed);\r
+> +    else\r
+> +        *out = talloc_realloc (ctx, *out, char, needed);\r
+> +\r
+> +    if (*out == NULL)\r
+> +        return 0;\r
+> +\r
+> +    *out_size = needed;\r
+> +    }\r
+> +    return 1;\r
+> +}\r
+> +\r
+> +hex_status_t\r
+> +hex_encode (void *ctx, const char *in, char **out, size_t *out_size)\r
+> +{\r
+> +\r
+> +    const unsigned char *p;\r
+\r
+The casts to unsigned char * below bother me. Perhaps this should be\r
+just const char *, with the only cast being in the sprintf?\r
+\r
+> +    char *q;\r
+> +\r
+> +    size_t escape_count = 0;\r
+> +    size_t len = 0;\r
+> +    size_t needed;\r
+> +\r
+> +    assert (ctx); assert (in); assert (out); assert (out_size);\r
+> +\r
+> +    for (p = (unsigned char *) in; *p; p++) {\r
+> +    escape_count += (!is_output (*p));\r
+> +    len++;\r
+> +    }\r
+> +\r
+> +    needed = len + escape_count * 2 + 1;\r
+\r
+I wonder if it would be clearer if escape_count and len were ditched,\r
+and the for loop just did:\r
+\r
+       needed += is_output (*p) ? 1 : 3;\r
+\r
+and another needed++ after the loop for NUL. And maybe s/needed/len/\r
+after that.\r
+\r
+> +\r
+> +    if (*out == NULL)\r
+> +    *out_size = 0;\r
+> +\r
+> +    if (!maybe_realloc (ctx, needed, out, out_size))\r
+> +    return HEX_OUT_OF_MEMORY;\r
+> +\r
+> +    q = *out;\r
+> +    p = (unsigned char *) in;\r
+> +\r
+> +    while (*p) {\r
+> +    if (is_output (*p)) {\r
+> +        *q++ = *p++;\r
+> +    } else {\r
+> +        sprintf (q, "%%%02x", *p++);\r
+> +        q += 3;\r
+> +    }\r
+> +    }\r
+> +\r
+> +    *q = '\0';\r
+> +    return HEX_SUCCESS;\r
+> +}\r
+> +\r
+> +/* Hex decode 'in' to 'out'.\r
+> + *\r
+> + * This must succeed for in == out to support hex_decode_inplace().\r
+> + */\r
+> +static hex_status_t\r
+> +hex_decode_internal (const char *in, unsigned char *out)\r
+> +{\r
+> +    char buf[3];\r
+> +\r
+> +    while (*in) {\r
+> +    if (*in == escape_char) {\r
+> +        char *endp;\r
+> +\r
+> +        /* This also handles unexpected end-of-string. */\r
+> +        if (!isxdigit ((unsigned char) in[1]) ||\r
+> +            !isxdigit ((unsigned char) in[2]))\r
+> +            return HEX_SYNTAX_ERROR;\r
+> +\r
+> +        buf[0] = in[1];\r
+> +        buf[1] = in[2];\r
+> +        buf[2] = '\0';\r
+> +\r
+> +        *out = strtoul (buf, &endp, 16);\r
+> +\r
+> +        if (endp != buf + 2)\r
+> +            return HEX_SYNTAX_ERROR;\r
+> +\r
+> +        in += 3;\r
+> +        out++;\r
+> +    } else {\r
+> +        *out++ = *in++;\r
+> +    }\r
+> +    }\r
+> +\r
+> +    *out = '\0';\r
+> +\r
+> +    return HEX_SUCCESS;\r
+> +}\r
+> +\r
+> +hex_status_t\r
+> +hex_decode_inplace (char *s)\r
+> +{\r
+> +    /* A decoded string is never longer than the encoded one, so it is\r
+> +     * safe to decode a string onto itself. */\r
+> +    return hex_decode_internal (s, (unsigned char *) s);\r
+> +}\r
+> +\r
+> +hex_status_t\r
+> +hex_decode (void *ctx, const char *in, char **out, size_t * out_size)\r
+> +{\r
+> +    const char *p;\r
+> +    size_t escape_count = 0;\r
+> +    size_t needed = 0;\r
+> +\r
+> +    assert (ctx); assert (in); assert (out); assert (out_size);\r
+> +\r
+> +    size_t len = strlen (in);\r
+> +\r
+> +    for (p = in; *p; p++)\r
+> +    escape_count += (*p == escape_char);\r
+> +\r
+> +    needed = len - escape_count * 2 + 1;\r
+\r
+Same as above for counting the needed size. It would also save scanning\r
+the input string twice (strlen and for loop).\r
+\r
+> +\r
+> +    if (!maybe_realloc (ctx, needed, out, out_size))\r
+> +    return HEX_OUT_OF_MEMORY;\r
+> +\r
+> +    return hex_decode_internal (in, (unsigned char *) *out);\r
+> +}\r
+> diff --git a/util/hex-escape.h b/util/hex-escape.h\r
+> new file mode 100644\r
+> index 0000000..5182042\r
+> --- /dev/null\r
+> +++ b/util/hex-escape.h\r
+> @@ -0,0 +1,41 @@\r
+> +#ifndef _HEX_ESCAPE_H\r
+> +#define _HEX_ESCAPE_H\r
+> +\r
+> +typedef enum hex_status {\r
+> +    HEX_SUCCESS = 0,\r
+> +    HEX_SYNTAX_ERROR,\r
+> +    HEX_OUT_OF_MEMORY\r
+> +} hex_status_t;\r
+> +\r
+> +/*\r
+> + * The API for hex_encode() and hex_decode() is modelled on that for\r
+> + * getline.\r
+> + *\r
+> + * If 'out' points to a NULL pointer a char array of the appropriate\r
+> + * size is allocated using talloc, and out_size is updated.\r
+> + *\r
+> + * If 'out' points to a non-NULL pointer, it assumed to describe an\r
+> + * existing char array, with the size given in *out_size.  This array\r
+> + * may be resized by talloc_realloc if needed; in this case *out_size\r
+> + * will also be updated.\r
+> + *\r
+> + * Note that it is an error to pass a NULL pointer for any parameter\r
+> + * of these routines.\r
+> + */\r
+> +\r
+> +hex_status_t\r
+> +hex_encode (void *talloc_ctx, const char *in, char **out,\r
+> +            size_t *out_size);\r
+> +\r
+> +hex_status_t\r
+> +hex_decode (void *talloc_ctx, const char *in, char **out,\r
+> +            size_t *out_size);\r
+> +\r
+> +/*\r
+> + * Non-allocating hex decode to decode 's' in-place. The length of the\r
+> + * result is always equal to or shorter than the length of the\r
+> + * original.\r
+> + */\r
+> +hex_status_t\r
+> +hex_decode_inplace (char *s);\r
+> +#endif\r
+> -- \r
+> 1.7.10.4\r
+>\r
+> _______________________________________________\r
+> notmuch mailing list\r
+> notmuch@notmuchmail.org\r
+> http://notmuchmail.org/mailman/listinfo/notmuch\r