Search-based regular expression inference on a GPU

Valizadeh, Mojtaba; Berger, Martin

3591274.pdf (1.44 MB)

Search-based regular expression inference on a GPU

conference contribution

posted on 2023-07-07, 09:47 authored by Mojtaba ValizadehMojtaba Valizadeh, Martin BergerMartin Berger

Regular expression inference (REI) is a supervised machine learning and program synthesis problem that takes a cost metric for regular expressions, and positive and negative examples of strings as input. It outputs a regular expression that is precise (i.e., accepts all positive and rejects all negative examples), and minimal w.r.t.To the cost metric. We present a novel algorithm for REI over arbitrary alphabets that is enumerative and trades off time for space. Our main algorithmic idea is to implement the search space of regular expressions succinctly as a contiguous matrix of bitvectors. Collectively, the bitvectors represent, as characteristic sequences, all sub-languages of the infix-closure of the union of positive and negative examples. Mathematically, this is a semiring of (a variant of) formal power series. Infix-closure enables bottom-up compositional construction of larger from smaller regular expressions using the operations of our semiring. This minimises data movement and data-dependent branching, hence maximises data-parallelism. In addition, the infix-closure remains unchanged during the search, hence search can be staged: first pre-compute various expensive operations, and then run the compute intensive search process. We provide two C++ implementations, one for general purpose CPUs and one for Nvidia GPUs (using CUDA). We benchmark both on Google Colab Pro: The GPU implementation is on average over 1000x faster than the CPU implementation on the hardest benchmarks.

History

Publication status

Published

File Version

Published version

Journal

Proceedings of the ACM on Programming Languages

ISSN

2475-1421

Publisher

Association for Computing Machinery (ACM)

Publisher URL

http://dx.doi.org/10.1145/3591274

External DOI

https://doi.org/10.1145/3591274

Volume

7

Page range

1317-1339

Event name

44th ACM SIGPLAN Conference on Programming Language Design and Implementation

Event location

Orlando, Florida, United States

Event type

Conference

Event date

Sat 17 - Wed 21 June 2023

Department affiliated with

Informatics Publications

Peer reviewed?

Yes

Usage metrics

Keywords

Search-Based Regular Expression Inference on a GPU

Licence

CC BY 4.0

Exports

RefWorks

BibTeX

Ref. manager

Endnote

DataCite

NLM

DC

Search-based regular expression inference on a GPU

History

Publication status

File Version

Journal

ISSN

Publisher

Publisher URL

External DOI

Volume

Page range

Event name

Event location

Event type

Event date

Department affiliated with

Peer reviewed?

Usage metrics

Categories

Keywords

Licence

Exports