From: Michal Nazarewicz Date: Mon, 17 Sep 2012 14:13:35 +0000 (+0200) Subject: Re: [PATCH v3 2/9] parse-time-string: add a date/time parser to notmuch X-Git-Url: http://git.tremily.us/gitweb.cgi?a=commitdiff_plain;h=c98e0a6f3c344a9d0f754f6844b01765b9e80e8a;p=notmuch-archives.git Re: [PATCH v3 2/9] parse-time-string: add a date/time parser to notmuch --- diff --git a/de/7a847d5b3a15e0566cd2273327edd7606056af b/de/7a847d5b3a15e0566cd2273327edd7606056af new file mode 100644 index 000000000..b103dfd9f --- /dev/null +++ b/de/7a847d5b3a15e0566cd2273327edd7606056af @@ -0,0 +1,1086 @@ +Return-Path: +X-Original-To: notmuch@notmuchmail.org +Delivered-To: notmuch@notmuchmail.org +Received: from localhost (localhost [127.0.0.1]) + by olra.theworths.org (Postfix) with ESMTP id 4831C431FAF + for ; Mon, 17 Sep 2012 07:13:51 -0700 (PDT) +X-Virus-Scanned: Debian amavisd-new at olra.theworths.org +X-Spam-Flag: NO +X-Spam-Score: -0.7 +X-Spam-Level: +X-Spam-Status: No, score=-0.7 tagged_above=-999 required=5 + tests=[DKIM_SIGNED=0.1, DKIM_VALID=-0.1, RCVD_IN_DNSWL_LOW=-0.7] + autolearn=disabled +Received: from olra.theworths.org ([127.0.0.1]) + by localhost (olra.theworths.org [127.0.0.1]) (amavisd-new, port 10024) + with ESMTP id Fw+F4127W2uw for ; + Mon, 17 Sep 2012 07:13:48 -0700 (PDT) +Received: from mail-bk0-f53.google.com (mail-bk0-f53.google.com + [209.85.214.53]) (using TLSv1 with cipher RC4-SHA (128/128 bits)) + (No client certificate requested) + by olra.theworths.org (Postfix) with ESMTPS id E5754431FAE + for ; Mon, 17 Sep 2012 07:13:47 -0700 (PDT) +Received: by bkwj4 with SMTP id j4so2779763bkw.26 + for ; Mon, 17 Sep 2012 07:13:45 -0700 (PDT) +DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; + s=20120113; h=sender:from:to:subject:in-reply-to:organization:references + :user-agent:x-face:face:x-pgp:x-pgp-fp:date:message-id:mime-version + :content-type; bh=SXKUsmwu/XziMg4brNjKTGDYW8B15ZqO0UtdkTxkpPg=; + b=pLAHGftO/0ZKngdQR98x5OteyD04pbkPE59nLF2fQ3h76SUlNmHX5sjZQJwyx6s2pF + ad35K6xDG+73CclcRlTNLsbFXiTJvW9hfGIPUugxCwnOTOqhdK2asJvo4roNfFl+quI6 + 22rGZnbPVbY4xk6s3EodUJb1fp7cp307dxCWlhQq+FIFLfBF5sKriaD3Vov0K17mjym0 + aXzird+Y149jm70qyev9OiWphi2TuSdkJNKX9FNqUBZE7IXwwXm6LqlBjbXNiweYQ5AL + RnxZiUGRnlaqEfNOVVK4PprlD61UTWBjO/LNnsOBvNVwk3qUpTmbHeual8eAX7QLq1PS /baw== +X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; + d=google.com; s=20120113; + h=sender:from:to:subject:in-reply-to:organization:references + :user-agent:x-face:face:x-pgp:x-pgp-fp:date:message-id:mime-version + :content-type:x-gm-message-state; + bh=SXKUsmwu/XziMg4brNjKTGDYW8B15ZqO0UtdkTxkpPg=; + b=kGbgTWuWqje8sC096jaqF5rLNOd+lXV5wVegczAmw9cTMtOtfnOwT8AOqOf55IfC9M + d47g8Cvg7TGhhXa2/v9uY0v4hfon224VVfp23p17xX/AKEqXqvmOLmRBVTgvvCHIjkYB + ZGJwFW9yADj0MyM3aDq7mm2heFzuzelu6RLLGBtVjqHuivJnxbjQDs5N6V+/UV/WykXa + ygRbJii0hjpoMpLvww7EXUFS7/hY6+G12XQ59rQyJ1Bb6yII5ri4CaESxlFjWrcNKhCD + fe/ZvnLsdsbYglRrM35OlkaYlQC4bpYTDdt7VsnOsTmKRbxCr7tqtyXWRAxowvcvKWv6 + YmLg== +Received: by 10.204.148.83 with SMTP id o19mr4351950bkv.74.1347891225206; + Mon, 17 Sep 2012 07:13:45 -0700 (PDT) +Received: by 10.204.148.83 with SMTP id o19mr4351935bkv.74.1347891224863; + Mon, 17 Sep 2012 07:13:44 -0700 (PDT) +Received: from mpn-glaptop ([172.28.91.186]) + by mx.google.com with ESMTPS id 14sm5629228bkw.15.2012.09.17.07.13.41 + (version=TLSv1/SSLv3 cipher=OTHER); + Mon, 17 Sep 2012 07:13:43 -0700 (PDT) +Sender: Michal Nazarewicz +From: Michal Nazarewicz +To: Jani Nikula , notmuch@notmuchmail.org, + David Bremner +Subject: Re: [PATCH v3 2/9] parse-time-string: add a date/time parser to + notmuch +In-Reply-To: <878vcerurc.fsf@nikula.org> +Organization: http://mina86.com/ +References: + <89741ec9a9687fca8b30aa1a4877392d355dd3ce.1347484177.git.jani@nikula.org> + <878vcerurc.fsf@nikula.org> +User-Agent: Notmuch/0.14+22~g8bdc16b (http://notmuchmail.org) Emacs/24.2.50.1 + (x86_64-unknown-linux-gnu) +X-Face: PbkBB1w#)bOqd`iCe"Ds{e+!C7`pkC9a|f)Qo^BMQvy\q5x3?vDQJeN(DS?|-^$uMti[3D*#^_Ts"pU$jBQLq~Ud6iNwAw_r_o_4]|JO?]}P_}Nc&"p#D(ZgUb4uCNPe7~a[DbPG0T~!&c.y$Ur,=N4RT>]dNpd; KFrfMCylc}gc??'U2j,!8%xdD +Face: iVBORw0KGgoAAAANSUhEUgAAADAAAAAwBAMAAAClLOS0AAAAJFBMVEWbfGlUPDDHgE57V0jUupKjgIObY0PLrom9mH4dFRK4gmjPs41MxjOgAAACQElEQVQ4jW3TMWvbQBQHcBk1xE6WyALX1069oZBMlq+ouUwpEQQ6uRjttkWP4CmBgGM0BQLBdPFZYPsyFUo6uEtKDQ7oy/U96XR2Ux8ehH/89Z6enqxBcS7Lg81jmSuujrfCZcLI/TYYvbGj+jbgFpHJ/bqQAUISj8iLyu4LuFHJTosxsucO4jSDNE0Hq3hwK/ceQ5sx97b8LcUDsILfk+ovHkOIsMbBfg43VuQ5Ln9YAGCkUdKJoXR9EclFBhixy3EGVz1K6eEkhxCAkeMMnqoAhAKwhoUJkDrCqvbecaYINlFKSRS1i12VKH1XpUd4qxL876EkMcDvHj3s5RBajHHMlA5iK32e0C7VgG0RlzFPvoYHZLRmAC0BmNcBruhkE0KsMsbEc62ZwUJDxWUdMsMhVqovoT96i/DnX/ASvz/6hbCabELLk/6FF/8PNpPCGqcZTGFcBhhAaZZDbQPaAB3+KrWWy2XgbYDNIinkdWAFcCpraDE/knwe5DBqGmgzESl1p2E4MWAz0VUPgYYzmfWb9yS4vCvgsxJriNTHoIBz5YteBvg+VGISQWUqhMiByPIPpygeDBE6elD973xWwKkEiHZAHKjhuPsFnBuArrzxtakRcISv+XMIPl4aGBUJm8Emk7qBYU8IlgNEIpiJhk/No24jHwkKTFHDWfPniR4iw5vJaw2nzSjfq2zffcE/GDjRC2dn0J0XwPAbDL84TvaFCJEU4Oml9pRyEUhR3Cl2t01AoEjRbs0sYugp14/4X5n4pU4EHHnMAAAAAElFTkSuQmCC +X-PGP: 50751FF4 +X-PGP-FP: AC1F 5F5C D418 88F8 CC84 5858 2060 4012 5075 1FF4 +Date: Mon, 17 Sep 2012 16:13:35 +0200 +Message-ID: +MIME-Version: 1.0 +Content-Type: multipart/mixed; boundary="=-=-=" +X-Gm-Message-State: ALoCoQlLV3rFEX4IG9wEvoVab+ezDzAM8zmuD17sqgf1ybc6TCmgyQFx8l1ChUNjRZeaVXQNmaS1v0n/3G+F109dmqrBUKdT2ejBkVH2TpXKvY1vBYySU+bW2mZeImXtGJVDJ3juu6Uq5D48AlAQ7SQdi6oFC+2j+WWA37XDrT1V/VW7z7cUhG4KBRF47sAZV7ucsT2Xe/uF +X-BeenThere: notmuch@notmuchmail.org +X-Mailman-Version: 2.1.13 +Precedence: list +List-Id: "Use and development of the notmuch mail system." + +List-Unsubscribe: , + +List-Archive: +List-Post: +List-Help: +List-Subscribe: , + +X-List-Received-Date: Mon, 17 Sep 2012 14:13:51 -0000 + +--=-=-= +Content-Type: text/plain; charset=utf-8 +Content-Transfer-Encoding: quoted-printable + +> On Thu, 13 Sep 2012, Michal Nazarewicz wrote: +>> Have you consider doing the same in bison? I consider the code totally +>> unreadable and unmaintainable. + +On Thu, Sep 13 2012, Jani Nikula wrote: +> I do not think you could easily do everything that this parser does in +> bison. But then I'm not an expert in bison, and I have zero ambition to +> become one. So I'm biased, and I'm open about it. + +Bison can do a lot of weird stuff including modifying how lexer +interpretes tokens even while parsing given grammar rule. + +> Even so, if you're suggesting doing this in bison would make this +> totally readable and maintainable, I urge you to have a good look at +> [1]. Note that it also does less in more lines of code. (And using it +> as-is in notmuch has pretty much been turned down in the past.) +> +> Finally, I also suggest you actually read and review the code, pointing +> out concrete issues in readability or maintainability that you +> see. Especially since an earlier version has received comment "[I]t +> looks very nice to me. It is pleasantly nice to read." [2]. What you're +> doing is worthless bikeshedding otherwise. + +I'm sorry. I sometime tend to go into extremes with my statements, so +yes, the =E2=80=9Ctotally unreadable=E2=80=9D was a over statement on my pa= +rt. + +My point was however that parsing is a solved problem, and for +non-trivial parsers one needs to ask herself whether it's worth trying +to implement the logic, or maybe using a parser generator is just +simpler. + +And in this particular case, my feeling is that bison is easier to read +and modify. + +To add some merit to my statement, I attach a bison parser. + +It supports ranges as so: + the specific moment with duration dependent + on specification. How duration is figured out + is described in the next paragraph. + .. dates >=3D and < , so for instance + =E2=80=9Cyesterday..0=E2=80=9D days yields results from yesterday. + .. dates < + .. dates >=3D + ++ a shorthand of =E2=80=9C.. + =E2=80=9D. + This is useful for things like: =E2=80=9C2012q1++2 + quarters=E2=80=9D which is equivalent to + =E2=80=9C2012/01/01..2012/07/01=E2=80=9D, ie. the first two + quarters of 2012. + +It supports specifications as: + '@' + Raw timestamp. It's duration is one second. + + (seconds | minutes | hours | days | weeks | fortnights) [ago] + moves the date by given number of units in the future or + in the past (if =E2=80=9Cago=E2=80=9D is given). can be preceded + by sign. + + This specification's duration is whatever unit was used, + ie. one second, one minute, one hour, one day, one week + or one fortnight. So =E2=80=9C7 days ago=E2=80=9D and =E2=80=9C1 week ag= +o=E2=80=9D + specify the same moment, but they hay different + durations. + + (months | quarters | years) [ago] + Like above, but calendar months are used which do not + always have the same length. If applying the offset + ends up with a day of the month out of range, the day + is capped to the last day of the month. + + yesterday + Moves one day back. [*] Note that because of [*] this + is not quivalent to =E2=80=9C-1day=E2=80=9D. + YYYY/MM/DD + YYYY-MM-DD + MM-DD-YYYY + DD Month YYYY + Month DD YYYY + Sets date accordingly. [*] =E2=80=9CMonth=E2=80=9D is a human readable + month name. + Month [DD] [YYYY] + If either day or year is missing, given component of the + date is not changed. [*] Also, if day is missing, the + duration is set to one month rather than one day (but + see caveats described in [*]). + YYYY q Q + Sets date to the beginning of quarter Q, ie. =E2=80=9C2012q2=E2=80=9D is + roughly the same as =E2=80=9C2012/04/01=E2=80=9D. [*] Sets duration to + three moths but see caveats described in [*]. + + midnight | noon + Sets time to 0:00:00 and 12:00:00 respectively. Has + duration of 1 hour. + HH:MM:SS [am | pm] + HH:MM [am | pm] + HH (am | pm) + Sets time accordingly with the part that is not + specified set to zero. Duration depends on how many + components are missing, ie. =E2=80=9CHH (am|pm)=E2=80=9D has a duration of + on hour, =E2=80=9CHH:MM=E2=80=9D has a duration of one minute and + =E2=80=9CHH:MM:SS=E2=80=9D has a duration of one second. + +[*] Formats specifying the date will zero the time to midnight unless + the time has already been specified (ie. =E2=80=9Cyesterday=E2=80=9D is= + roughly the + same as =E2=80=9Cyesterday midnight=E2=80=9D, but =E2=80=9Cnoon yesterd= +ay=E2=80=9D still keeps time + as noon. + + Also, if the time has not been specified, those formats will set + duration to one day (with two exception), so =E2=80=9Cyesterday=E2=80= +=9D has + a duration of one day, but =E2=80=9Cyesterday midnight=E2=80=9D, even t= +hough it + specifies the same moment's beginning, has a duration of one hour. + +Purposly, I have not added support for MM/DD/YYYY or DD/MM/YYYY as well +as two-digit years. I feel this would only add confusion. + +--- + .gitignore | 3 + + Makefile | 17 ++ + date-parser-grammar.y | 173 ++++++++++++++++++ + date-parser.c | 476 +++++++++++++++++++++++++++++++++++++++++++++= +++++ + date-parser.h | 59 ++++++ + test.c | 44 +++++ + 6 files changed, 772 insertions(+), 0 deletions(-) + create mode 100644 .gitignore + create mode 100644 Makefile + create mode 100644 date-parser-grammar.y + create mode 100644 date-parser.c + create mode 100644 date-parser.h + create mode 100644 test.c + +diff --git a/.gitignore b/.gitignore +new file mode 100644 +index 0000000..b73c782 +--- /dev/null ++++ b/.gitignore +@@ -0,0 +1,3 @@ ++test ++*.o ++date-parser-grammar.tab.* +diff --git a/Makefile b/Makefile +new file mode 100644 +index 0000000..8d95f71 +--- /dev/null ++++ b/Makefile +@@ -0,0 +1,17 @@ ++CFLAGS +=3D -std=3Dc99 -Wextra -Werror -pedantic ++ ++test: test.o date-parser.o date-parser-grammar.tab.o ++test.o: test.c date-parser.h ++date-parser.o: date-parser.c date-parser.h ++date-parser.o: date-parser-grammar.tab.h ++ ++date-parser-grammar.tab.c: date-parser-grammar.y ++ bison $< ++ ++date-parser-grammar.tab.h: date-parser-grammar.tab.c ++date-parser-grammar.tab.o: date-parser-grammar.tab.c date-parser.h ++date-parser-grammar.tab.o: CPPFLAGS +=3D -Wno-unreachable-code ++ ++clean: ++ rm -f date-parser-grammar.output date-parser-grammar.tab.* \ ++ *.o test +diff --git a/date-parser-grammar.y b/date-parser-grammar.y +new file mode 100644 +index 0000000..38ddfca +--- /dev/null ++++ b/date-parser-grammar.y +@@ -0,0 +1,173 @@ ++/* Date parser bison grammar file ++ * Copyright (c) 2012 Google Inc. ++ * Written by Michal Nazarewicz ++ * ++ * This program is free software: you can redistribute it and/or modify ++ * it under the terms of the GNU General Public License as published by ++ * the Free Software Foundation, either version 3 of the License, or ++ * (at your option) any later version. ++ * ++ * This program is distributed in the hope that it will be useful, ++ * but WITHOUT ANY WARRANTY; without even the implied warranty of ++ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the ++ * GNU General Public License for more details. ++ * ++ * You should have received a copy of the GNU General Public License ++ * along with this program. If not, see http://www.gnu.org/licenses/ . */ ++ ++%code requires { ++ ++#ifndef YYSTYPE ++# define YYSTYPE long ++#endif ++#ifndef YYLTYPE ++# define YYLTYPE struct yylocation ++#endif ++ ++#ifdef YYLLOC_DEFAULT ++# undef YYLLOC_DEFAULT ++#endif ++#define YYLLOC_DEFAULT(Cur, Rhs, N) do { \ ++ if (N) { \ ++ (Cur).start =3D YYRHSLOC(Rhs, 1).start; \ ++ (Cur).end =3D YYRHSLOC(Rhs, N).end; \ ++ } else { \ ++ (Cur) =3D YYRHSLOC(Rhs, 0); \ ++ } \ ++} while (0) ++} ++ ++%code{ ++#include "date-parser-grammar.tab.h" ++#include "date-parser.h" ++ ++#define ASSERT(cond, loc, message) do { \ ++ if (!(cond)) { \ ++ parse_date_print_error(&loc, message); \ ++ YYERROR; \ ++ } \ ++} while (0) ++} ++ ++%locations ++%defines ++%error-verbose ++%define api.pure ++ ++%parse-param {struct date *ret} ++%parse-param {const char **inputp} ++%lex-param {const char **inputp} ++ ++%token T_NUMBER "" /* Always positive. */ ++%token T_NUMBER_4 "####" /* Four digit number. */ ++%token T_AGO "ago" ++/* Also used for minutes, hours, days and weeks. */ ++%token T_SECONDS "seconds" ++/* Also used for quarters and years */ ++%token T_MONTHS "months" ++%token T_YESTERDAY "yesterday" ++%token T_AMPM "am/pm" ++%token T_HOUR "" ++%token T_MONTH "" ++ ++%expect 3 /* Two shift/reduce conflicts caused by year_maybe, and onde ++ * caused by day_maybe. */ ++ ++%% ++ /* For backwards compatibility, just a number and nothing else ++ * is treated as timestamp */ ++input : number { date_set_from_stamp(ret, $1) } ++ | date ++ ; ++ ++date : part ++ | date part ++ ; ++ ++part : integer "seconds" ago_maybe { ++ ASSERT(date_add_seconds(ret, $1 * $3, $2), @$, ++ "offset ends up in date out of range") ++ } ++ | integer "months" ago_maybe { ++ ASSERT(date_add_months(ret, $1 * $3, $2), @$, ++ "offset ends up in date out of range") ++ } ++ | "yesterday" { ++ ASSERT(date_set_yesterday(ret), @$, ++ "offset ends up in date out of range") ++ } ++ ++ | '@' number { date_set_from_stamp(ret, $2) } ++ | "" { date_set_time(ret, $1, -1, -1, -1) } ++ ++ /* HH:MM, HH:MM:SS, HH:MM am/pm, HH:MM:SS am/pm */ ++ | number ':' number seconds_maybe ampm_maybe { ++ ASSERT(date_set_time(ret, $1, $3, $4, $5), @$, ++ "invalid time") ++ } ++ ++ | number "am/pm" { /* HH am/pm */ ++ ASSERT(date_set_time(ret, $1, -1, -1, $2), @$, "invalid hour") ++ } ++ ++ | "####" '/' "" '/' "" { /* YYYY/MM/DD */ ++ ASSERT(date_set_date(ret, $1, $3, $5), @$, "invalid date") ++ } ++ | "####" '-' "" '-' "" { /* YYYY-MM-DD */ ++ ASSERT(date_set_date(ret, $1, $3, $5), @$, "invalid date") ++ } ++ | "" '-' "" '-' "####" { /* DD-MM-YYYY */ ++ ASSERT(date_set_date(ret, $5, $3, $1), @$, "invalid date") ++ } ++ /* No MM/DD/YYYY or DD/MM/YYYY because it's confusing. */ ++ ++ | "" "" year_maybe { /* 1 January 2012 */ ++ ASSERT(date_set_date(ret, $3, $2, $1), @$, "invalid date") ++ } ++ | "" day_maybe year_maybe { /* January 1 2012 */ ++ ASSERT(date_set_date(ret, $3, $1, $2), @$, "invalid date") ++ } ++ ++ | "####" 'q' "" { /* Quarter, 2012q1 */ ++ ASSERT(date_set_quarter(ret, $1, $3), @$, "invalid quarter"); ++ } ++ ; ++ ++number : "" { $$ =3D $1 } ++ | "####" { $$ =3D $1 } ++ ; ++ ++integer : number { $$ =3D $1 } ++ | '-' number { $$ =3D -$2 } ++ ; ++ ++ago_maybe ++ : /* empty */ { $$ =3D 1 } ++ | "ago" { $$ =3D -1 } ++ ; ++ ++seconds_maybe ++ : /* empty */ { $$ =3D -1 } ++ | ':' "" { $$ =3D $2 } ++ ; ++ ++ampm_maybe ++ : /* empty */ { $$ =3D -1 } ++ | "am/pm" { $$ =3D $$ } ++ /* For people who like writing "a.m." or "p.m." and since dot ++ * is ignored by the lexer (ie. it's treated just like white ++ * space), dot is lost. */ ++ | 'a' 'm' { $$ =3D 0 } ++ | 'p' 'm' { $$ =3D 0 } ++ ; ++ ++day_maybe ++ : /* empty */ { $$ =3D -1 } ++ | "" { $$ =3D $1 } ++ ; ++ ++year_maybe ++ : /* empty */ { $$ =3D -1 } ++ | "####" { $$ =3D $1 } ++ ; ++%% +diff --git a/date-parser.c b/date-parser.c +new file mode 100644 +index 0000000..c1701bd +--- /dev/null ++++ b/date-parser.c +@@ -0,0 +1,476 @@ ++/* Date parser implementation file. ++ * Copyright (c) 2012 Google Inc. ++ * Written by Michal Nazarewicz ++ * ++ * This program is free software: you can redistribute it and/or modify ++ * it under the terms of the GNU General Public License as published by ++ * the Free Software Foundation, either version 3 of the License, or ++ * (at your option) any later version. ++ * ++ * This program is distributed in the hope that it will be useful, ++ * but WITHOUT ANY WARRANTY; without even the implied warranty of ++ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the ++ * GNU General Public License for more details. ++ * ++ * You should have received a copy of the GNU General Public License ++ * along with this program. If not, see http://www.gnu.org/licenses/ . */ ++ ++#define _POSIX_C_SOURCE 1 ++ ++#include "date-parser.h" ++#include "date-parser-grammar.tab.h" ++ ++#include ++#include ++#include ++#include ++#include ++ ++ ++static bool is_valid_year(int year) { ++ /* TODO: Get the actual time_t range. */ ++ return year >=3D 1970 && year < 2037; ++} ++ ++ ++/***************************** Basic date helpers ************************= +***/ ++ ++static int days_in_months[] =3D { ++ 31, 28, 31, 30, 31, 30, 31, 31, 30, 31, 30, 31 ++}; ++ ++static inline int is_leap(int year) { ++ return year % 4 =3D=3D 0 && (year % 100 || year % 400 =3D=3D 0); ++} ++ ++static inline int days_in_month(int year, int month) { ++ return days_in_months[month] + (month =3D=3D 1 ? is_leap(year) : 0); ++} ++ ++static inline int days_in_year(int year) { ++ return 365 + is_leap(year); ++} ++ ++static inline int min(int a, int b) { ++ return a < b ? a : b; ++} ++ ++ ++/****************************** Date manipulation ************************= +***/ ++ ++struct date { ++ struct tm tm; ++ int dur_sec, dur_mon, has_time; ++}; ++ ++ ++void date_set_from_stamp(struct date *ret, long stamp) { ++ time_t t =3D stamp; ++ localtime_r(&t, &ret->tm); ++ ret->dur_sec =3D 1; ++ ret->dur_mon =3D 0; ++ ret->has_time =3D 1; ++} ++ ++static void date_set_to_now(struct date *ret) { ++ time_t t =3D time(NULL); ++ localtime_r(&t, &ret->tm); ++ ret->dur_sec =3D 1; ++ ret->dur_mon =3D 0; ++ ret->has_time =3D 0; ++} ++ ++static void date_zero_time(struct date *ret, int dur_sec, int dur_mon) { ++ if (!ret->has_time) { ++ ret->tm.tm_hour =3D ret->tm.tm_min =3D ret->tm.tm_sec =3D 0; ++ ret->dur_sec =3D dur_sec; ++ ret->dur_mon =3D dur_mon; ++ } ++} ++ ++bool date_add_seconds(struct date *ret, long num, long unit) { ++ time_t t =3D mktime(&ret->tm) + num * unit; ++ localtime_r(&t, &ret->tm); ++ ret->dur_sec =3D unit; ++ ret->dur_mon =3D 0; ++ return true; /* TODO add validation */ ++} ++ ++bool date_set_yesterday(struct date *ret) { ++ if (ret->tm.tm_mday !=3D 1) { ++ --ret->tm.tm_mday; ++ } else if (ret->tm.tm_mon) { ++ --ret->tm.tm_mon; ++ ret->tm.tm_mday =3D days_in_month(ret->tm.tm_year + 1900, ++ ret->tm.tm_mon); ++ } else if (is_valid_year(1900 + ret->tm.tm_year - 1)) { ++ --ret->tm.tm_year; ++ ret->tm.tm_mon =3D 11; ++ ret->tm.tm_mday =3D 31; ++ } else { ++ return false; ++ } ++ date_zero_time(ret, 24 * 3600, 0); ++ ret->tm.tm_isdst =3D -1; ++ return true; ++} ++ ++bool date_add_months(struct date *ret, long num, long unit) { ++ long y; ++ ++ y =3D ret->tm.tm_year + 1900; ++ num =3D num * unit + ret->tm.tm_mon; ++ if (num < 0) { ++ y -=3D -num / 12; ++ num =3D 11 - (-num % 12); ++ } else { ++ y +=3D num / 12; ++ num %=3D 12; ++ } ++ if (!is_valid_year(y)) { ++ return false; ++ } ++ ++ ret->tm.tm_year =3D y - 1900; ++ ret->tm.tm_mon =3D num; ++ ret->tm.tm_mday =3D min(ret->tm.tm_mday, ++ days_in_month(ret->tm.tm_year + 1900, ++ ret->tm.tm_mon)); ++ if (!ret->has_time) { ++ ret->dur_sec =3D 0; ++ ret->dur_mon =3D unit; ++ } ++ ret->tm.tm_isdst =3D -1; ++ return true; ++} ++ ++bool date_set_time(struct date *ret, long h, long m, long s, int ampm) { ++ if (m > 60 || s > 60 || h > 23) { ++ return false; ++ } ++ ++ if (ampm !=3D -1) { ++ if (!h || h > 12) { ++ return false; ++ } ++ if (h !=3D 12) { ++ h +=3D ampm * 12; ++ } else if (ampm) { /* 12 pm */ ++ h =3D 12; ++ } else { ++ /* 12 am is 0 the next day, so adjust date */ ++ date_add_seconds(ret, 1, 24 * 3600); ++ h =3D 0; ++ } ++ } ++ ++ if (m =3D=3D -1) { ++ ret->dur_sec =3D 3600; ++ m =3D s =3D 0; ++ } else if (s =3D=3D -1) { ++ ret->dur_sec =3D 60; ++ s =3D 0; ++ } else { ++ ret->dur_sec =3D 1; ++ } ++ ret->dur_mon =3D 0; ++ ++ ret->tm.tm_hour =3D h; ++ ret->tm.tm_min =3D m; ++ ret->tm.tm_sec =3D s; ++ ret->tm.tm_isdst =3D -1; ++ ++ ret->has_time =3D 1; ++ return true; ++} ++ ++bool date_set_date(struct date *ret, long y, long m, long d) { ++ if (y =3D=3D -1) { ++ y =3D ret->tm.tm_year + 1900; ++ } else if (!is_valid_year(y)) { ++ return false; ++ } ++ if (m < 1 || m > 12 || ++ (d !=3D -1 && (d < 1 || d > days_in_month(y, m)))) { ++ return false; ++ } ++ ret->tm.tm_year =3D y - 1900; ++ ret->tm.tm_mon =3D m - 1; ++ ret->tm.tm_mday =3D d =3D=3D -1 ? 1 : d; ++ if (d =3D=3D -1) { ++ date_zero_time(ret, 0, 1); ++ } else { ++ date_zero_time(ret, 24 * 3600, 0); ++ } ++ ret->tm.tm_isdst =3D -1; ++ return true; ++} ++ ++bool date_set_quarter(struct date *ret, long y, long q) { ++ if (!is_valid_year(y) || q < 1 || q > 4) { ++ return false; ++ } ++ ret->tm.tm_year =3D y - 1900; ++ ret->tm.tm_mon =3D (q - 1) * 3; ++ ret->tm.tm_mday =3D 1; ++ date_zero_time(ret, 0, 3); ++ ret->tm.tm_isdst =3D -1; ++ return true; ++} ++ ++ ++#define TOKEN(str, token, num) { str, sizeof(str) - 1, token, num } ++#define ABBR(len) { NULL, len, 0, 0 } ++ ++static struct token { ++ const char *str; ++ size_t len; ++ int token; ++ long num; ++} tokens_array[] =3D { ++ TOKEN("ago", T_AGO, 0), ++ ++ TOKEN("am", T_AMPM, 0), ++ TOKEN("pm", T_AMPM, 1), ++ ++ TOKEN("seconds", T_SECONDS, 1), ABBR(3), ABBR(6), ++ TOKEN("minutes", T_SECONDS, 60), ABBR(3), ABBR(6), ++ TOKEN("hours", T_SECONDS, 3600), ABBR(4), ABBR(1), ++ TOKEN("days", T_SECONDS, 24 * 3600), ABBR(3), ABBR(1), ++ TOKEN("weeks", T_SECONDS, 7 * 24 * 3600), ABBR(4), ++ TOKEN("fortnights", T_SECONDS, 14 * 24 * 3600), ABBR(9), ++ ++ TOKEN("months", T_MONTHS, 1), ABBR(5), ++ TOKEN("quarters", T_MONTHS, 3), ABBR(7), ++ TOKEN("years", T_MONTHS, 12), ABBR(4), ++ ++ TOKEN("yesterday", T_YESTERDAY, 0), ++ ++ TOKEN("midnight", T_HOUR, 0), ++ TOKEN("noon", T_HOUR, 12), ++ ++ TOKEN("january", T_MONTH, 1), ABBR(3), ++ TOKEN("february", T_MONTH, 2), ABBR(3), ++ TOKEN("march", T_MONTH, 3), ABBR(3), ++ TOKEN("april", T_MONTH, 4), ABBR(3), ++ TOKEN("may", T_MONTH, 5), ++ TOKEN("june", T_MONTH, 6), ABBR(3), ++ TOKEN("july", T_MONTH, 7), ABBR(3), ++ TOKEN("august", T_MONTH, 8), ABBR(3), ++ TOKEN("september", T_MONTH, 9), ABBR(4), ABBR(3), ++ TOKEN("october", T_MONTH, 10), ABBR(3), ++ TOKEN("november", T_MONTH, 11), ABBR(3), ++ TOKEN("december", T_MONTH, 12), ABBR(3), ++ ++ { NULL, 0, 0, 0 }, ++}; ++ ++#undef TOKEN ++#undef ABBR ++ ++ ++static struct token locale_tokens_array[2*12 + 1]; ++static bool locale_tokens_populated =3D false; ++ ++static void populate_locale_tokens(void) { ++ static const char *mon_formats[] =3D { "%b", "%B" }; ++ static char locale_buffer[1024]; ++ ++ char *buf =3D locale_buffer, *end =3D buf + sizeof(locale_buffer); ++ struct token *out =3D locale_tokens_array; ++ struct tm tm; ++ unsigned i; ++ ++ tm.tm_sec =3D 0; ++ tm.tm_min =3D 0; ++ tm.tm_hour =3D 0; ++ tm.tm_mday =3D 10; ++ tm.tm_year =3D 100; ++ tm.tm_isdst =3D 0; ++ ++ for (tm.tm_mon =3D 0; tm.tm_mon < 12; ++tm.tm_mon) { ++ for (i =3D 0; i < 2; ++i) { ++ out->len =3D strftime(buf, end - buf, ++ mon_formats[i], &tm); ++ if (!out->len) { ++ continue; ++ } ++ out->str =3D buf; ++ buf +=3D out->len; ++ out->token =3D T_MONTH; ++ out->num =3D tm.tm_mon + 1; ++ ++out; ++ } ++ } ++ ++ out->len =3D 0; ++} ++ ++static const struct token *find_token(const struct token *tk, ++ const char *str, size_t len) { ++ const struct token *ret; ++ ++ for (; tk->len; ++tk) { ++ if (tk->str) { ++ ret =3D tk; ++ } ++ if (tk->len =3D=3D len && !strncasecmp(str, ret->str, len)) { ++ return ret; ++ } ++ } ++ ++ return NULL; ++} ++ ++ ++/* Treat '_' and '.' as white space so that people don't have to quote ++ * the argument when specifying it on command line. */ ++#define SKIP_WHITE_SPACE(ch) do { \ ++ while (isspace(*ch) || *ch =3D=3D '_' || *ch =3D=3D '.') { \ ++ ++ch; \ ++ } \ ++} while (0) ++ ++ ++int yylex(YYSTYPE *valp, struct yylocation *loc, const char **inputp) { ++ const char *ch =3D *inputp, *str; ++ const struct token *tk; ++ ++ SKIP_WHITE_SPACE(ch); ++ ++ /* End of data */ ++ if (*ch =3D=3D 0) { ++ return EOF; ++ } ++ ++ loc->start =3D ch; ++ ++ /* Parse number */ ++ if (isdigit(*ch)) { ++ errno =3D 0; ++ *valp =3D strtol(ch, (char**)&ch, 10); ++ ++ loc->end =3D ch; ++ *inputp =3D ch; ++ ++ if (errno) { ++ parse_date_print_error(loc, "number out of range"); ++ return 256; ++ } ++ return ch - loc->start =3D=3D 4 ? T_NUMBER_4 : T_NUMBER; ++ } ++ ++ if (!isalpha(*ch)) { ++ *inputp =3D ch + 1; ++ loc->end =3D ch + 1; ++ return *ch; ++ } ++ ++ /* So it's a string token. */ ++ str =3D ch; ++ while (isalpha(*ch)) { ++ ++ch; ++ } ++ loc->end =3D ch; ++ *inputp =3D ch; ++ ++ tk =3D find_token(tokens_array, str, ch - str); ++ if (tk) { ++ *valp =3D tk->num; ++ return tk->token; ++ } ++ ++ /* Let's try with locale strings. */ ++ if (!locale_tokens_populated) { ++ populate_locale_tokens(); ++ locale_tokens_populated =3D true; ++ } ++ tk =3D find_token(locale_tokens_array, str, ch - str); ++ if (tk) { ++ *valp =3D tk->num; ++ return tk->token; ++ } ++ ++ /* If it's just one letter, return it converted to lower case */ ++ if (ch - str =3D=3D 1) { ++ return tolower(*str); ++ } ++ ++ parse_date_print_error(loc, "unrecognised token"); ++ return 256; ++} ++ ++ ++/**************************** Parsing interface **************************= +***/ ++ ++static bool parse_date(struct date *ret, const char *from, char *to) { ++ char tmp; ++ int res; ++ if (to) { ++ tmp =3D *to; ++ *to =3D '\0'; ++ } ++ res =3D yyparse(ret, &from); ++ if (to) { ++ *to =3D tmp; ++ } ++ return res =3D=3D 0; ++} ++ ++bool parse_range(char *arg, time_t *from, time_t *to) { ++ bool left =3D false, right =3D false; ++ struct date a, b; ++ char *dd, *pp; ++ ++ SKIP_WHITE_SPACE(arg); ++ if (!*arg) { ++ fprintf(stderr, "empty range argument\n"); ++ return false; ++ } ++ ++ dd =3D strstr(arg, ".."); ++ pp =3D strstr(arg, "++"); ++ if ((dd && pp) || ++ (dd && strstr(dd + 2, "..")) || ++ (pp && strstr(pp + 2, "++"))) { ++ fprintf(stderr, ++ "%s: at most one of '..' or '++' can be used\n", arg); ++ return false; ++ } ++ ++ if (dd || pp) { ++ char *ch =3D dd ? dd : pp; ++ left =3D ch !=3D arg; ++ SKIP_WHITE_SPACE(ch); ++ right =3D *ch; ++ } ++ if (pp && (!right || !left)) { ++ fprintf(stderr, ++ "%s: '++' requires expression on both sides\n", arg); ++ return false; ++ } ++ ++ if (left || !right) { /* date:.. or date: */ ++ date_set_to_now(&a); ++ if (!parse_date(&a, arg, dd ? dd : pp)) { ++ return false; ++ } ++ } ++ if (right) { /* date:.. or date:.. */ ++ if (pp) { ++ b =3D a; ++ } else { ++ date_set_to_now(&b); ++ } ++ if (!parse_date(&b, (dd ? dd : pp) + 2, NULL)) { ++ return false; ++ } ++ } else if (!left) { /* date:date */ ++ left =3D right =3D true; /* convert to date:.. */ ++ b =3D a; ++ if ((b.dur_sec && !date_add_seconds(&b, b.dur_sec, 1)) || ++ (b.dur_mon && !date_add_months(&b, b.dur_mon, 1))) { ++ right =3D false; /* convert to date:.. */ ++ } ++ } ++ ++ *from =3D left ? mktime(&a.tm) : 0; ++ *to =3D right ? mktime(&b.tm) : (time_t)((unsigned long)~(time_t)0 >> 1); ++ return true; ++} +diff --git a/date-parser.h b/date-parser.h +new file mode 100644 +index 0000000..fb0f19b +--- /dev/null ++++ b/date-parser.h +@@ -0,0 +1,59 @@ ++/* Date parser header file. ++ * Copyright (c) 2012 Google Inc. ++ * Written by Michal Nazarewicz ++ * ++ * This program is free software: you can redistribute it and/or modify ++ * it under the terms of the GNU General Public License as published by ++ * the Free Software Foundation, either version 3 of the License, or ++ * (at your option) any later version. ++ * ++ * This program is distributed in the hope that it will be useful, ++ * but WITHOUT ANY WARRANTY; without even the implied warranty of ++ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the ++ * GNU General Public License for more details. ++ * ++ * You should have received a copy of the GNU General Public License ++ * along with this program. If not, see http://www.gnu.org/licenses/ . */ ++ ++#ifndef H_DATE_PARSER_H ++#define H_DATE_PARSER_H ++ ++#include ++#include ++#include ++ ++bool parse_range(char *arg, time_t *from, time_t *to); ++ ++/* For parser */ ++struct date; ++ ++struct yylocation { ++ const char *start, *end; ++}; ++ ++static inline void parse_date_print_error(struct yylocation *loc, ++ const char *message) { ++ fprintf(stderr, "%.*s: %s\n", ++ (int)(loc->end - loc->start), loc->start, message); ++} ++ ++static inline int yyerror(struct yylocation *loc, struct date *ret, ++ const char **inputp, const char *message) { ++ ret =3D ret; /* make compiler happy */ ++ inputp =3D inputp; ++ parse_date_print_error(loc, message); ++ return 0; ++} ++ ++int yylex(long *valp, struct yylocation *loc, const char **inputp); ++int yyparse(struct date *ret, const char **inputp); ++ ++void date_set_from_stamp(struct date *ret, long stamp); ++bool date_add_seconds(struct date *ret, long num, long unit); ++bool date_set_yesterday(struct date *ret); ++bool date_add_months(struct date *ret, long num, long unit); ++bool date_set_time(struct date *ret, long h, long m, long s, int ampm); ++bool date_set_date(struct date *ret, long y, long m, long d); ++bool date_set_quarter(struct date *ret, long y, long q); ++ ++#endif +diff --git a/test.c b/test.c +new file mode 100644 +index 0000000..c4e2d9c +--- /dev/null ++++ b/test.c +@@ -0,0 +1,44 @@ ++/* Date parser testing application. ++ * Copyright (c) 2012 Google Inc. ++ * Written by Michal Nazarewicz ++ * ++ * This program is free software: you can redistribute it and/or modify ++ * it under the terms of the GNU General Public License as published by ++ * the Free Software Foundation, either version 3 of the License, or ++ * (at your option) any later version. ++ * ++ * This program is distributed in the hope that it will be useful, ++ * but WITHOUT ANY WARRANTY; without even the implied warranty of ++ * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the ++ * GNU General Public License for more details. ++ * ++ * You should have received a copy of the GNU General Public License ++ * along with this program. If not, see http://www.gnu.org/licenses/ . */ ++ ++#define _POSIX_C_SOURCE 1 ++ ++#include ++#include ++#include ++ ++#include "date-parser.h" ++ ++int main(void) { ++ char buf[1024], *ch; ++ time_t from, to; ++ struct tm tm; ++ ++ setlocale(LC_ALL, ""); ++ ++ while (fgets(buf, sizeof buf, stdin)) { ++ if (parse_range(buf, &from, &to)) { ++ localtime_r(&from, &tm); ++ ch =3D buf + strftime(buf, sizeof buf / 2, ++ "[%Y/%m/%d %H:%M:%S %Z, ", &tm); ++ localtime_r(&to, &tm); ++ strftime(ch, sizeof buf / 2, ++ "%Y/%m/%d %H:%M:%S %Z)\n", &tm); ++ fputs(buf, stdout); ++ } ++ } ++} +--=20 +1.7.7.3 + +--=20 +Best regards, _ _ +.o. | Liege of Serenely Enlightened Majesty of o' \,=3D./ `o +..o | Computer Science, Micha=C5=82 =E2=80=9Cmina86=E2=80=9D Nazarewicz = + (o o) +ooo +------------------ooO--(_)--Ooo-- +--=-=-= +Content-Type: multipart/signed; boundary="==-=-="; + micalg=pgp-sha1; protocol="application/pgp-signature" + +--==-=-= +Content-Type: text/plain + + +--==-=-= +Content-Type: application/pgp-signature + +-----BEGIN PGP SIGNATURE----- +Version: GnuPG v1.4.10 (GNU/Linux) + +iQIcBAEBAgAGBQJQVzAPAAoJECBgQBJQdR/0e8UP/jA6s6ANgAHVEcfvd4nvCpaZ +DEBV1hpm6HGJDQ/P1NPRafn6k4ZdaZokCK1qWMjgmYNOQ6jmfNLhieGuEXXaTsQs +NRmroRQfWpDAIeQgsF/9dIAZgFoCigLKt/ge0p4e/rHBePmYOfxr3T9RBuGw9drv +/zZ+fpv55P6iZP+xzl8TxeeVL9CDbzkuytnH3Kg1nFKgEFHlRoZxjeq/iELKXz/n +reMF3cxoAZkYXwW5WntA/mWUqUFptWgBYZSi6bPhg6KEbJdr2lu6L3kxrYbuSUn9 +MBRIqG60OWVwguoCUbHE71mv39tnINb7uHTZumvdQkBlr2deeXPzho1xCwqMQH17 +sZ6IkSe2hhfaCxgBWhddzvHHr7TPIuAVZGHWucJuWOHCvBasn21ARtcOn/q0dhVT +XGTvAw072lbo2l/ceaEQTGDrj1IV8Xo/1OeQU4y6n56hrVDknb4zwkdlk1BIcNu/ +Upi2vgewDETFl+NuLE7rr2/4UNcXAgBQIZnwlL5+vo/EV2/eTsTY0TemYxq/ppZC +/YtV1x35sc1pN5LswSGU6tDeFp0+TNXaPXEW48IPrE5OhmaCrQT84DK62+UJDJ8r +Ga4MmQh6GFQ3EemKoDHim59gp2hC3HOrHiQ9MWlf2GG0zYVOKto02qjPCjoeRjzc +ZTefvm76BBIEpTfwa1fx +=STYS +-----END PGP SIGNATURE----- +--==-=-=-- + +--=-=-=--