The Shape of Money Laundering: Subgraph Representation Learning on the Blockchain with the Elliptic2 Dataset

  • 2024-05-01 05:55:30
  • Claudio Bellei, Muhua Xu, Ross Phillips, Tom Robinson, Mark Weber, Tim Kaler, Charles E. Leiserson, Arvind, Jie Chen
  • 0

Abstract

Subgraph representation learning is a technique for analyzing localstructures (or shapes) within complex networks. Enabled by recent developmentsin scalable Graph Neural Networks (GNNs), this approach encodes relationalinformation at a subgroup level (multiple connected nodes) rather than at anode level of abstraction. We posit that certain domain applications, such asanti-money laundering (AML), are inherently subgraph problems and mainstreamgraph techniques have been operating at a suboptimal level of abstraction. Thisis due in part to the scarcity of annotated datasets of real-world size andcomplexity, as well as the lack of software tools for managing subgraph GNNworkflows at scale. To enable work in fundamental algorithms as well as domainapplications in AML and beyond, we introduce Elliptic2, a large graph datasetcontaining 122K labeled subgraphs of Bitcoin clusters within a background graphconsisting of 49M node clusters and 196M edge transactions. The datasetprovides subgraphs known to be linked to illicit activity for learning the setof "shapes" that money laundering exhibits in cryptocurrency and accuratelyclassifying new criminal activity. Along with the dataset we share our graphtechniques, software tooling, promising early experimental results, and newdomain insights already gleaned from this approach. Taken together, we findimmediate practical value in this approach and the potential for a new standardin anti-money laundering and forensic analytics in cryptocurrencies and otherfinancial networks.

 

Quick Read (beta)

loading the full paper ...