terraform

Terraformを使ってAWSにWebアプリケーションの実行環境を立てる (EC2立てるまで)

投稿日 2021年07月25日

Webアプリケーション実行環境をIaCで管理したい.
Terraformでクラウド構成を作ってAnsibleでミドルウェアをインストールしたい.
BeanstalkやLightsailのようなPaaSではなくTerraformを使ってVPCから自前で作ってみる.
この記事はEC2を立てるまでが範囲.
次の記事でAnsibleを使って立てたEC2にミドルウェアをインストールする.

【目次】

∨この記事で紹介する範囲
∨Terraformの導入
∨git secretsの導入
∨ディレクトリ構成
∨tfstateの保存先の定義
∨credentialsの書き方
∨providerの定義
∨エントリポイント
∨VPCモジュール
∨EC2モジュール
∨実行
∨疎通確認

この記事で紹介する範囲

この記事ではTerraformを使ってAWS上に以下の構成を作るまでを書いてみる.

とはいえTerraformの習得が8割くらいのモチベなので実用性はあまり重視しない.
サブネットをプライベートとパブリックに分けてみたい.
プライベートにDB(MySQL), パブリックにWebサーバ(nginx).
ひとまずALBは配置しない.

Terraformの導入

Ansibleもそうだけれども, アプリを保守している期間って割と長いもので、
その間, 構成管理ツール側のバージョンが上がってしまう傾向がある.
そうすぐに古い書き方が使えなくなることはないが, 警告が出まくって気分がよくない.
構成管理ツールの古いバージョンを残しておきたい, どのバージョンを使うか選びたい, という期待がある.
rbenvやpyenvのようにTerraform自体のバージョンを管理するtfenvをインストールしておき,
この記事を書いた日の最新である 1.0.3 をインストールすることにする.


$ brew install tfenv
$ tfenv --version
tfenv 2.2.2

$ tfenv list-remote
.1.0-alpha20210714
1.1.0-alpha20210630
1.1.0-alpha20210616
1.0.3
1.0.2
1
...
$  tfenv install 1.0.3
...

$ tfenv list
1.0.3

$ tfenv use 1.0.3
Switching default version to v1.0.3
Switching completed

$ terraform version
Terraform v1.0.3
on darwin_amd64

git secretsの導入

AWSのcredentialsなどを誤ってcommitしてしまう事故を防ぐためにgit secretsを導入する.
commit時に内容を検証してくれて, もしそれらしきファイルがあればリジェクトしてくれる.
どこまで見てくれるのか未検証だけれども入れておく.
Laravelの.env_staging等に書いたcredentialsがどう扱われるか後で検証する.


$ brew install git-secrets
$ git secrets --install
✓ Installed commit-msg hook to .git/hooks/commit-msg
✓ Installed pre-commit hook to .git/hooks/pre-commit
✓ Installed prepare-commit-msg hook to .git/hooks/prepare-commit-msg
$ git secrets --register-aws
OK

ディレクトリ構成

勉強用の小さな環境を作るのだけれども, 今後の拡張性については考慮しておきたい.
割と規定されている傾向があるAnsibleと比較して,Terraformは自由な印象.
以下の記事を参考にさせて頂きました.
Terraformなにもわからないけどディレクトリ構成の実例を晒して人類に貢献したい


iac
├── dev
│   ├── backend.tf
│   ├── main.tf -> ../shared/main.tf
│   ├── provider.tf -> ../shared/provider.tf
│   ├── versions.tf -> ../shared/versions.tf
│   ├── terraform.tfvars
│   └── variables.tf -> ../shared/variables.tf
└── shared
    ├── main.tf
    ├── provider.tf
    ├── variables.tf
    └── modules
        ├── vpc
        │   ├── eip.tf
        │   ├── internet_gateway.tf
        │   ├── nat_gateway.tf
        │   ├── routetables.tf
        │   ├── subnet.tf
        │   ├── vpc.tf
        │   ├── outputs.tf
        │   └── variables.tf
        └── ec2
            ├── ec2.tf
            ├── keypair.tf
            ├── network_interface.tf
            ├── security_group.tf
            ├── outputs.tf
            └── variables.tf

tfstateの保存先の定義

tfstate は Terraformが管理しているリソースの現在の状態を表すファイル.
terraformは「リソースを記述したファイル」と「現在の状態」の差分を埋めるように処理を行うが,
いちいち「現在の状態」を調べにいくとパフォーマンスが悪化するため, ファイルに保存される.
(確かにAnsibleは毎回「現在の状態」を調べにいっているっぽく,これが結構遅くて毎回イライラする)

デフォルトだとローカルに作られるが, それだとチーム開発で共有できないので,
S3等に作るのが良くあるパターン.
Terraformでは”バックエンド”という概念で扱われる. “バックエンド”を以下のように記述する.

バックエンドの定義はterraformの前段にあり, S3 bucketとDynamoDB tableを手動で作っておく必要がある.
変数を使うことができないのでハードコードしないといけない. 議論があるらしい.
key,secretを書く代わりにprofileを書くことで, 構成管理可能になる.
(同じprofile名をチームで共有しないといけない…)

backendをS3にする際にS3のbucketをどう作るか問題はいろいろ議論があるようで,
いずれ以下の記事を参考にしてよしなにbucketを作れるようにしたい.

dynamodb_tableを設定すると、そこにロックファイルを作ってくれるようになる.
多人数で同じ構成管理を触るときに便利.
Backend の S3 や DynamoDB 自体を terraform で管理するセットアップ方法


terraform {
  backend "s3" {
    region = "ap-northeast-1"
    profile = "ikuty"
    bucket = "terraform-state-dev"
    key    = "terraform-state-dev.tfstate"
    dynamodb_table = "terraform-state-lock-dev"
  }
}

credentialsの書き方

ルートにある terraform.tfvarsというファイルを置いておくと、
そこに記述した内容を変数に注入することができる.
“注入”という言葉で良いのか不明だが、定義した変数の初期値を設定してくれる.

credentialsを構成管理に登録するのはご法度.
terraform.tfvarsを構成管理外として何らかの方法で環境にコピーする.
多くのツールで採用されている「よくあるパターン」.

他に,applyコマンドに直接渡したり, 環境変数で指定したりできるが,
Terraform公式は.tfvarsを推奨している.


aws_access_key_id = "AKI*****************"
aws_secret_access_key = "9wc*************************************"
aws_region = "ap-northeast-1"

providerの定義

プロバイダとは, 要は”AWS”,”Azure”,”GCP”.. のような粒度の何か.
Terraformは結構な種類のプロバイダに対応していて「どのプロバイダを使うか」を定義する.
今回はAWSを使う. dev.tfvarsに記述しておいたCredentialsを変数で受けて設定する.

以下,変数の定義方法, デフォルト値の設定方法を示している.
.tfvarsに記述した同名の変数について,terraformが値を設定してくれる.


variable "aws_access_key_id" {}
variable "aws_secret_access_key" {}
variable "aws_region" {
    default = "ap-northeast-1"
}

provider "aws" {
    access_key = "${var.aws_access_key_id}"
    secret_key = "${var.aws_secret_access_key}"
    region = "${var.aws_region}"
}

エントリポイント

Terraformのエントリポイントはルートに置いた”main.tf”.
ディレクトリ構成を凝らないのであれば、main.tf に全てをベタ書きすることもできる.
今回、devやstg, prod のような環境ごとにルートを分ける構成を作りたいのだが、
main.tf 自体は環境ごとに差異が無いことを前提にしている.

./shared/main.tf というファイルを作成し、
各環境ごとの main.tf を ./shared/main.tf の Symbolic Link とする.

main.tf でリソースの定義はおこなわない. 同階層の./modules にモジュール定義があるが,
main.tf は ./modules以下の各モジュールに変数を渡すだけ.

VPCの作成とEC2の作成を各モジュールに分割した.
各モジュールのOutputのスコープはモジュールまでなので、
例えばEC2モジュールからVPCモジュールのVPC IDを直接受け取れない.
main.tf はモジュールの上に位置するため、このようにモジュール間で変数を共有できる.


module "vpc" {
        source = "../shared/modules/vpc"
}
module "ec2" {
        source = "../shared/modules/ec2"
        vpc_id = module.vpc.myVPC.id
        private_subnet_id = module.vpc.private_subnet.id
        public_subnet_id = module.vpc.public_subnet.id
}

VPCモジュール

./shared/modules/vpc以下にVPCモジュールを構成するファイルを配置する.

スコープがVPCモジュールに閉じたローカル変数を定義する.
以下のようにしておくと、モジュール内から local.vpc_cidr.dev のように値を取得できる.


locals {
        vpc_cidr = {
                dev = "10.1.0.0/16"
        }
        subnet_cidr = {
                private = "10.1.2.0/24"
                public = "10.1.1.0/24"
        }
}

VPCを1個作る. VPCのCIDRは10.1.0.0/16.


resource "aws_vpc" "myVPC" {
        cidr_block = local.vpc_cidr.dev
        instance_tenancy = "default"
        enable_dns_support = "true"
        enable_dns_hostnames = "false"
        tags = {
                Name = "myVPC"
        }
}

作ったVPC内にサブネットを2個作る. 1つはPrivate用. もう1つはPublic用.
PrivateサブネットのCIDRは10.1.2.0/24. PublicサブネットのCIDRは10.1.1.0/24.
AZは両方同じで “ap-northeast-1a”.
map_public_ip_on_launchをtrueとしておくと,
そこで立ち上げたEC2に自動的にpublic ipが振られる.


resource "aws_subnet" "public_1a" {
        vpc_id = aws_vpc.myVPC.id
        depends_on = [aws_vpc.myVPC]
        availability_zone = "ap-northeast-1a"
        cidr_block = local.subnet_cidr.public
        map_public_ip_on_launch = true
        tags = {
                Name = "public-1a"
        }
}
resource "aws_subnet" "private_1a" {
        vpc_id = aws_vpc.myVPC.id
        depends_on = [aws_vpc.myVPC]
        availability_zone = "ap-northeast-1a"
        cidr_block = local.subnet_cidr.private
        tags = {
                Name = "private-1a"
        }
}

VPCに紐づくInternet Gatewayを作る.


resource "aws_internet_gateway" "myGW" {
        vpc_id = "${aws_vpc.myVPC.id}"
        depends_on = [aws_vpc.myVPC]
        tags = {
                Name = "my Internet Gateway"
        }
}

Privateサブネットからインターネットに繋ぐために、
PublicサブネットにNAT Gatewayを作りたい.
NAT Gateway用のEIPを作る.


resource "aws_eip" "nat_gateway" {
        vpc = true
        depends_on = [aws_internet_gateway.myGW]
        tags = {
                Name = "Eip for Nat gateway"
        }
}

PublicサブネットにNAT Gatewayを作る.
EIPは上で作成したものを使う.


resource "aws_nat_gateway" "myNatGW" {
        allocation_id = aws_eip.nat_gateway.id
        subnet_id = aws_subnet.public_1a.id
        depends_on = [aws_internet_gateway.myGW]
        tags = {
                Name = "my Nat Gateway"
        }
}

ルートテーブル. いろいろなところで書かれていた内容を試してようやく動くものができた.
VPCにはデフォルトで「メインルートテーブル」が作られる.
メインルートテーブルはいじっていない.

以下、Private, Publicサブネットそれぞれのためのルートテーブルを定義している.
PublicサブネットからInternet Gatewayに繋ぐ. PrivateサブネットからNAT Gatewayに繋ぐ.


# Route table for public
# public
resource "aws_route_table" "public" {
        vpc_id = aws_vpc.myVPC.id
        depends_on = [aws_internet_gateway.myGW]
        tags = {
                Name = "my Route Table for public"
        }
}
# private
resource "aws_route_table" "private" {
        vpc_id = aws_vpc.myVPC.id
        depends_on = [aws_internet_gateway.myGW]
        tags = {
                Name = "my Route Table for private"
        }
}

# Route table association
# public
resource "aws_route_table_association" "public" {
        subnet_id = aws_subnet.public_1a.id
        route_table_id = aws_route_table.public.id
}
# private
resource "aws_route_table_association" "private" {
        subnet_id = aws_subnet.private_1a.id
        route_table_id = aws_route_table.private.id
}

# Routing for public
resource "aws_route" "public" {
        route_table_id = aws_route_table.public.id
        gateway_id = aws_internet_gateway.myGW.id
        destination_cidr_block = "0.0.0.0/0"
}

# Routing for private
resource "aws_route" "private" {
        route_table_id = aws_route_table.private.id
        gateway_id = aws_nat_gateway.myNatGW.id
        destination_cidr_block = "0.0.0.0/0"
}

EC2モジュール

./shared/modules/ec2以下にEC2モジュールを構成するファイルを配置する.

スコープがEC2モジュールに閉じたローカル変数を定義する.
main.tfからVPCモジュールのOutputをEC2モジュールに渡す必要があるが、
渡すデータを受けるためにEC2モジュール側で変数を定義しておく必要がある.


locals {
        private = {
                ip = "10.1.2.5"
                ami = "ami-0df99b3a8349462c6"
                instance_type = "t2.micro"
        }
        public = {
                ip = "10.1.1.5"
                ami = "ami-0df99b3a8349462c6"
                instance_type = "t2.micro"
        }
}
variable "vpc_id" {
        type = string
}
variable "private_subnet_id" {
        type = string
}
variable "public_subnet_id" {
        type = string
}

EC2にアクセスするための鍵ペア.
既に鍵ペアを持っているものとし、その公開鍵を渡す.
以下のようにすると、HostからSSHの-iオプションで秘密鍵を指定して接続できるようになる.


resource "aws_key_pair" "deployer" {
        key_name = "deployer"
        public_key = "{公開鍵}"
}

EC2に設定するセキュリティグループを作る.
この記事では, Private, Publicともに、インバウンドをSSHのみとした.
次の記事でPublicにHTTPを通す.
アウトバウンドとして全て通すようにしないとインスタンスから外にアクセスできなくなる(ハマった).


# Security group
resource "aws_security_group" "web_server_sg" {
        name = "web_server"
        description = "Allow http and https traffic."
        vpc_id = var.vpc_id
}

# Security group rule SSH(22)
resource "aws_security_group_rule" "web_inbound_ssh" {
        type = "ingress"
        from_port = 22
        to_port = 22
        protocol = "tcp"
        cidr_blocks = ["0.0.0.0/0"]
        security_group_id = aws_security_group.web_server_sg.id
}
resource "aws_security_group_rule" "web_outbound" {
        type = "egress"
        from_port = 0
        to_port = 0
        protocol = "-1"
        cidr_blocks = ["0.0.0.0/0"]
        ipv6_cidr_blocks = ["::/0"]
        security_group_id = aws_security_group.web_server_sg.id
}

# Security group
resource "aws_security_group" "db_server_sg" {
        name = "db_server"
        description = "Allow MySQL traffic."
        vpc_id = var.vpc_id
}

# Security group rule SSH(22)
resource "aws_security_group_rule" "db_inbound_ssh" {
        type = "ingress"
        from_port = 22
        to_port = 22
        protocol = "tcp"
        cidr_blocks = ["0.0.0.0/0"]
        security_group_id = aws_security_group.db_server_sg.id
}
resource "aws_security_group_rule" "db_outbound" {
        type = "egress"
        from_port = 0
        to_port = 0
        protocol = "-1"
        cidr_blocks = ["0.0.0.0/0"]
        ipv6_cidr_blocks = ["::/0"]
        security_group_id = aws_security_group.db_server_sg.id
}

ネットワークインターフェース.
セキュリティグループはEC2インスタンスではなくネットワークインターフェースに紐づく.
EC2(aws_instance)のsecurity_groupsに書けなくてハマった.


# public
resource "aws_network_interface" "public_1a" {
        subnet_id   = var.public_subnet_id
        private_ips = [local.public.ip]
        security_groups = [
                aws_security_group.web_server_sg.id
        ]
        tags = {
                Name = "public_subnet_network_interface"
        }
}
# private
resource "aws_network_interface" "private_1a" {
        subnet_id   = var.private_subnet_id
        private_ips = [local.private.ip]
        security_groups = [
                aws_security_group.db_server_sg.id
        ]
        tags = {
                Name = "private_subnet_network_interface"
        }
}

最後にEC2.


# Web Server
resource "aws_instance" "public" {
        ami = local.public.ami
        instance_type = local.public.instance_type
        key_name = aws_key_pair.deployer.id
        network_interface {
                network_interface_id = aws_network_interface.public_1a.id
                device_index = 0
        }
        credit_specification {
                cpu_credits = "unlimited"
        }
        root_block_device {
                volume_size = 20
                volume_type = "gp2"
                delete_on_termination = true
                tags = {
                        Name = "web-ebs"
                }
        }
        tags = {
                Name = "Web"
        }
}

# DB Server
resource "aws_instance" "private" {
        ami = local.private.ami
        instance_type = local.private.instance_type
        key_name = aws_key_pair.deployer.id
        network_interface {
                network_interface_id = aws_network_interface.private_1a.id
                device_index = 0
        }
        credit_specification {
                cpu_credits = "unlimited"
        }
        root_block_device {
                volume_size = 20
                volume_type = "gp2"
                delete_on_termination = true
                tags = {
                        Name = "db-ebs"
                }
        }
        tags = {
                Name = "DB"
        }
}

実行

作った.tfファイルを再生して環境を構築する.
validateでデバッグして、大体できたらplan(DryRun)で変更が正しそうか確認してみた.
が、評価しなければわからないものについてはDryRunではわからず、
結局applyが途中で止まって解決しないといけない.

ansibleと異なり冪等性が言われていなくて、applyで間違った構成を作ってしまうと、
その先、その構成を修正したとしても上手くいかないことがある.


$ cd "/path/to/dev"
$ terraform validate
Success! The configuration is valid.
$ terraform plan
...
$ terraform apply
...

出来たとして、Publicに立ったEC2のパブリックIPv4をメモる.

疎通確認

Host ->(SSH)-> Web ->(SSH)-> DB を試す. DBから外に繋がるか試す.
SSH Agent Forwardを使うと、Web EC2に秘密鍵を置かないで済む.
Web側のssh configにForwardAgent yesを指定しておく.


Host db
        HostName 10.1.2.5
        User ubuntu
        ForwardAgent yes

いざ.


$ ssh-add "{秘密鍵のパス}"
$ ssh -A ubuntu@{WebのパブリックIPv4}
ubuntu@ip-10-1-1-5$ ssh db
ubuntu@ip-10-1-2-5$ ping yahoo.co.jp
64 bytes from f1.top.vip.kks.yahoo.co.jp (183.79.135.206): icmp_seq=1 ttl=33 time=14.9 ms
64 bytes from f1.top.vip.kks.yahoo.co.jp (183.79.135.206): icmp_seq=2 ttl=33 time=14.6 ms
..

できた..

外部テーブルと使い方

使ったことがない機能のドキュメントを読んで詳しくなるシリーズ。ステージに配置したファイルに対してロードせずに直接クエリを実行できる仕組み。 [arst_toc tag=\"h4\"] 外部テーブルの基本ファイルの構造を自動的に解釈してテーブルにしてくれる訳ではなく、いったんレコードがVARIANT型のVALUEカラムに書かれる。外部テーブルに対するSELECT文の中で、VALUEカラム内の値を取得して新規に列を作る。この列を仮想列と言ったりする。外部テーブルは、クエリ実行のたびにファイルにアクセスすることになるため遅い。 Materialized Viewを作成することで、高速に外部テーブルにアクセスできる。複数ファイルとパーティション化外部テーブルの実体は外部ステージ上のファイルであって、通常、複数のファイルから構成される。複数のファイルが1つの外部テーブルとして扱われるところがポイント。 stackoverflowにドンピシャの記事があってとても参考になった。 Snowflake External Table Partition - Granular Path /appname/lob/ という論理ディレクトリの下に Granular にファイルが配置されている例。例えば、/appname/lob/2020/07/24/hoge.txt のように論理ディレクトリ以下に日付を使ってファイルが配置されることを想定している。 CREATE EXTERNAL TABLEする際、LOCATION を /appname/lob/ とすることで、 /appname/lob/以下に配置された複数のファイルが1つの外部テーブルで扱われる。その際、それぞれのファイル名が metadata$filename に渡される。公式によると、外部テーブルのパーティショニングを行うことが推奨されている。パーティション化された外部テーブル下記の例では、ファイル名からYYYY/MM/DDを抽出して日付化して、パーティションキーとしている。 CREATE OR REPLACE EXTERNAL TABLE Database.Schema.ExternalTableName( date_part date as to_date(substr(metadata$filename, 14, 10), \'YYYY/MM/DD\'), col1 varchar AS (value:col1::varchar)), col2 varchar AS (value:col2::varchar)) PARTITION BY (date_part) INTEGRATION = \'YourIntegration\' LOCATION=@SFSTG/appname/lob/ AUTO_REFRESH = true FILE_FORMAT = (TYPE = JSON); パーティショニングとは、名前が似ているがマイクロパーティションの事ではないと思う。概念的に似ていて、データを別々の塊として格納し、検索時にプルーニング的な何かを期待する。ガバナンス系オブジェクトとの関係以前、列レベルセキュリティのマスキングポリシーをまとめたように、テーブルと独立してマスキングポリシーを定義し、カラムに適用(Apply)することで列を保護する。 [clink url=\"https://ikuty.com/2023/03/31/column-level-security/\"] 外部テーブルは上記のように、VALUE列にKey-Valueで列値が入り、その後、SELECT文内で仮想列を定義する、という仕様で、外部テーブルに列がある訳ではない。このため、マスキングポリシーを適用する先がない、という問題が起こる。ただ、VALUE列全体にApplyすることは出来る様子。外部テーブル-セキュリティカラムイントロ Masking Policyはビューのカラムに対して適用することができるが、これは、外部テーブルに設定したMaterialized Viewにも設定できる。本記事の図のように、Materialized Viewを用意した上であれば、 Masking Policyを適用することができる。

クエリ同時実行特性の詳細な制御

なんだか分かった気がして分からないSnowflakeのウェアハウスの同時実行特性。ウェアハウスが1度に何個のクエリを同時処理するかだいぶ抽象化されているがウェアハウスは並列化された計算可能なコンピュータ。処理すべきクエリが大量にあった場合、1つのウェアハウスがそれらを同時に処理しようとする。 1個のウェアハウスが同時に最大何個のクエリを同時処理するか、 MAX_CONCURRENCY_LEVEL により制御できる。デフォルト値は 8 に設定されている。 MAX_CONCURRENCY_LEVEL を超えると、クエリはキューに入れられる。または、マルチクラスタウェアハウスの場合は、次のウェアハウスた立ち上がる。もちろん、クエリ1個が軽ければ同時処理は早く終わるが、重ければ時間がかかる。逆に言うと、MAX_CONCURRENCY_LEVEL を小さくすると、1つのウェアハウスが同時処理する最大クエリ数の上限が下がり、クエリ1個に割りあたるコンピューティングリソースが増える。 MAX_CONCURRENCY_LEVEL を大きくすると、クエリ1個に割り当たるリソースが減る。同時実行数を減らすとキューイングされる数が増える MAX_CONCURRENCY_LEVEL を小さくするとクエリの処理性能が上がるが、キューに入るクエリの数が増える。ちなみにデフォルトだと、キューに入ったクエリは永遠にキューに残り続ける。キューに入ったクエリをシステム側でキャンセルする時間を STATEMENT_QUEUED_TIMEOUT_IN_SECONDS により設定できる。ウェアハウスサイズとの関係ウェアハウスサイズアップはスケールアップ、マルチクラスタはスケールアウト、なんかと、テキトーに丸めて表現されたりするが、ウェアハウスサイズアップにより、コンピューティングリソースが増加するため、同じ MAX_CONCURRENCY_LEVEL であってもクエリ1個に割り当たるリソースが増えることで、結果的に所要時間が小さくなる。見かけ上、並列性が上がった風になる。ウェアハウスサイズ、ウェアハウス数の2つのファクタによりリソースの総量が決まるため、これらを固定した状態で MAX_CONCURRENCY_LEVEL を増やしたとしても、結果として処理できる量が増えたりはしない様子。同時実行性能が関係するパフォーマンスのベスプラお金をかけずに性能が上がるということはない。 MAX_CONCURRENCY_LEVELはあくまで微調整用途。大量の小さなクエリをガンガン処理するといったケースでは、ベスプラでは順当にウェアハウス数を増やして、その上でウェアハウスサイズを増やすべきとされている。この辺りが、「気分でウェアハウスサイズを増やしても性能上がらないなー」という現象の一因（、かもしれない）。

Golang + Gin カスタムバリデーション

Golang+GinによるAPI構築で使いそうなフィーチャーを試してみるシリーズ。今回はカスタムバリデーションを試してみる。 [clink implicit=\"false\" url=\"https://gin-gonic.com/ja/docs/examples/custom-validators/\" imgurl=\"https://gin-gonic.com/_astro/gin.D6H2T_2v_ZD2G7l.webp\" title=\"カスタムバリデーション\" excerpt=\"カスタムしたバリデーションを使用することもできます。サンプルコードも見てみてください。\"] [arst_toc tag=\"h4\"] ルーティングバリデーションを外部に移譲することで、ハンドラからロジック以外の冗長な処理を除くことができる。 Ginはカスタムバリデータを用意している。以下の例では、ユーザ登録を行うPOSTリクエストの例。組み込みのバリデーション・バインディングと合わせて、パスワードバリデーションロジックの追加を行っている。 package main import ( \"github.com/gin-gonic/gin\" \"github.com/gin-gonic/gin/binding\" \"github.com/go-playground/validator/v10\" \"github.com/ikuty/golang-gin/handlers\" ) func main() { // Ginエンジンの初期化 r := gin.Default() // カスタムバリデーターを登録 if v, ok := binding.Validator.Engine().(*validator.Validate); ok { handlers.InitCustomValidators(v) } // 7. カスタムバリデーション r.POST(\"/api/register\", handlers.RegisterValidatorHandler) // サーバー起動 r.Run(\":8080\") } ハンドラリクエストで受けたJSONをRegisterRequest構造体にバインディングする際に、組み込みのバリデーションルールを定義するのとは別に、strongpassword というカスタムルールを定義している。 strongpasswordルールの実体は strongPassword() 。例に出現するオブジェクトの使い方は、まぁこう使うのかぐらいで、ありがちな感じ。カスタムバリデータ関数がチェック結果をTrue/Falseで返せばよさそう。組み込みバリデータ、または、カスタムバリデータのバリデーション結果と文字列の対応を定義し、その文字列をレスポンスに付与して返す、というのは良くあるパターンで、 Ginで実装する場合は、また、カスタムバリデータのバリデーション結果と文字列の対応を定義しレスポンスに含める、というパターンは良くありそうで、構造体へのバインディングで発生したエラー(err)を取得し、 errに対する型アサーションを行った上で、errを validator.ValidationErrors型として扱う。動的型付けだと、発生したerrが本当に期待したオブジェクトなのか実行するまで分からなが、全ての処理が静的型付けを通して、実行前に実行可能であることが確認される。 package handlers import ( \"net/http\" \"regexp\" \"github.com/gin-gonic/gin\" \"github.com/go-playground/validator/v10\" ) // RegisterRequest はユーザー登録リクエストの構造体（高度なバリデーション付き） type RegisterRequest struct { Username string `json:\"username\" binding:\"required,min=3,max=20,alphanum\"` Email string `json:\"email\" binding:\"required,email\"` Password string `json:\"password\" binding:\"required,min=8,max=50,strongpassword\"` Age int `json:\"age\" binding:\"required,gte=18,lte=100\"` Website string `json:\"website\" binding:\"omitempty,url\"` Phone string `json:\"phone\" binding:\"omitempty,e164\"` // E.164 形式の電話番号 } // カスタムバリデーター: 強力なパスワードチェック func strongPassword(fl validator.FieldLevel) bool { password := fl.Field().String() // 最低1つの大文字、1つの小文字、1つの数字を含む hasUpper := regexp.MustCompile(`[A-Z]`).MatchString(password) hasLower := regexp.MustCompile(`[a-z]`).MatchString(password) hasNumber := regexp.MustCompile(`[0-9]`).MatchString(password) return hasUpper && hasLower && hasNumber } // RegisterValidatorHandler はカスタムバリデーターを使用するハンドラー func RegisterValidatorHandler(c *gin.Context) { var req RegisterRequest // JSON をバインド if err := c.ShouldBindJSON(&req); err != nil { // バリデーションエラーを詳細に返す c.JSON(http.StatusBadRequest, gin.H{ \"error\": \"Validation failed\", \"details\": formatValidationError(err), }) return } c.JSON(http.StatusCreated, gin.H{ \"message\": \"Registration successful\", \"username\": req.Username, \"email\": req.Email, }) } // formatValidationError はバリデーションエラーをわかりやすく整形 func formatValidationError(err error) []string { var errors []string if validationErrors, ok := err.(validator.ValidationErrors); ok { for _, e := range validationErrors { var message string switch e.Tag() { case \"required\": message = e.Field() + \" is required\" case \"email\": message = e.Field() + \" must be a valid email address\" case \"min\": message = e.Field() + \" must be at least \" + e.Param() + \" characters\" case \"max\": message = e.Field() + \" must be at most \" + e.Param() + \" characters\" case \"alphanum\": message = e.Field() + \" must contain only letters and numbers\" case \"gte\": message = e.Field() + \" must be greater than or equal to \" + e.Param() case \"lte\": message = e.Field() + \" must be less than or equal to \" + e.Param() case \"url\": message = e.Field() + \" must be a valid URL\" case \"e164\": message = e.Field() + \" must be a valid phone number (E.164 format)\" case \"strongpassword\": message = e.Field() + \" must contain at least one uppercase letter, one lowercase letter, and one number\" default: message = e.Field() + \" is invalid\" } errors = append(errors, message) } } else { errors = append(errors, err.Error()) } return errors } // InitCustomValidators はカスタムバリデーターを登録する func InitCustomValidators(v *validator.Validate) { v.RegisterValidation(\"strongpassword\", strongPassword) } 実行結果リクエストに対してバリデーションが行われ、期待通りバリデーションエラーがアサートされていて、アサートと対応するカスタム文字列がレスポンスに含まれていることが確認できる。 $ curl -X POST http://localhost:8080/api/register -H \"Content-Type: application/json\" -d \'{\"username\":\"john123\",\"email\":\"john@example.com\",\"password\":\"SecurePass123\",\"age\":25,\"website\":\"https://example.com\"}\' {\"email\":\"john@example.com\",\"message\":\"Registration successful\",\"username\":\"john123\"} 2. ユーザー名が短すぎる {\"details\":[\"Username must be at least 3 characters\"],\"error\":\"Validation failed\"} 3. 弱いパスワード（カスタムバリデーター） {\"details\":[\"Password must contain at least one uppercase letter, one lowercase letter, and one number\"],\"error\":\"Validation failed\"} 4. 年齢が18歳未満 {\"details\":[\"Age must be greater than or equal to 18\"],\"error\":\"Validation failed\"} まとめ組み込みバリデーションの他に、カスタムバリデーションを追加できることを確認した。静的型付けにより実行時エラーに頼ることのないある種の堅牢さがあることも見てとれた。

Fellegi-Sunterモデルに基づく確率的名寄せパッケージ Splinkを試してみる

[mathjax] Record Linkage、Entity Recognitionなど、いわゆる「名寄せ」問題には、割とさまざまな解決策が提案されている。その1つに確率論的な数学的背景を持つFellegi-Sunterモデルを実装したSplinkがある。数学的背景について咀嚼して説明してくれているサイトはあまりなく、Splinkの公式ドキュメントが唯一に近いソースだろうと思う。今回は5万人の人物データを使って重複を除去するSplinkの主要な機能のサンプルを追ってみた。数式を読んだところで実りは薄い、ということもあり、なるべく数式を読まずに読み進めたが、モデルのパラメタ、予測結果を解釈する際にどうしても数式を読まないとわからない箇所があった。特にベイズ統計について理解がかなり浅くて、正直なところ理解しきれてはいない。 Splinkの公式ドキュメントを読み進めていった際に気付いたメモを書いてみることにした。 [arst_toc tag=\"h4\"] 環境構築・準備環境構築の実装部分は主な学習要素ではないのでClaude Codeでサクッと作ってみた。まずはPlan modeで計画を立ててもらう。これぐらいだとほぼ言い値で使える。 DuckDB、Splinkを使う最小構成の環境を作ります。環境はDocker composeで構築します。計画を立ててください。 ⏺ DuckDB、Splinkを使うDocker Compose環境の計画を完了しました。計画概要 1. 設計文書: .tmp/design.mdで要件、アーキテクチャ、技術仕様を定義 2. タスク管理: .tmp/task.mdで詳細な実装手順を3フェーズに分けて整理主要コンポーネント - Jupyter Lab環境: DuckDB、Splink、必要なライブラリがプリインストール - データボリューム: サンプルデータ、notebook、出力用のディレクトリ - 最小構成: 開発・学習用途に特化した軽量環境次のステップ実装を開始する場合は、Phase 1から順番に進めます： 1. Docker設定ファイル作成 2. ディレクトリ構造構築 3. サンプル実装 4. テストと検証計画の詳細は.tmp/design.mdと.tmp/task.mdをご確認ください。実装を開始しますか？実際にファイルを作ってもらうと以下のような構成となる。得られる中間成果物である docker-compose.yml、Dockerfile は本記事では省略する。 . ├── data ├── docker-compose.yml ├── Dockerfile ├── notebooks ├── output └── requirements.txt 普通に http://localhost:8888 で JupyterLab が開く。使用するサンプルデータ 5万人の人物データを使って名寄せを行うサンプル。おそらくSplinkの用途として最初に思いつくやつ。 Splinkにデータをロードする前に必要なデータクリーニング手順について説明がある。公式によると、まずは行に一意のIDを割り当てる必要がある。データセット内で一意となるIDであって、重複除去した後のエンティティを識別するIDのことではない。 [clink implicit=\"false\" url=\"https://moj-analytical-services.github.io/splink/demos/tutorials/01_Prerequisites.html\" imgurl=\"https://user-images.githubusercontent.com/7570107/85285114-3969ac00-b488-11ea-88ff-5fca1b34af1f.png\" title=\"Data Prerequisites\" excerpt=\"Splink では、リンクする前にデータをクリーンアップし、行に一意の ID を割り当てる必要があります。このセクションでは、Splink にデータをロードする前に必要な追加のデータクリーニング手順について説明します。\"] 使用するサンプルデータは以下の通り。 from splink import splink_datasets df = splink_datasets.historical_50k df.head() データの分布を可視化 splink.exploratoryのprofile_columnsを使って分布を可視化してみる。 from splink import DuckDBAPI from splink.exploratory import profile_columns db_api = DuckDBAPI() profile_columns(df, db_api, column_expressions=[\"first_name\", \"substr(surname,1,2)\"]) 同じ姓・名の人が大量にいることがわかる。ブロッキングとブロッキングルールの評価テーブル内のレコードが他のレコードと「同一かどうか」を調べるためには、基本的には、他のすべてのレコードとの何らかの比較操作を行うこととなる。全てのレコードについて全てのカラム同士を比較したいのなら、対象のテーブルをCROSS JOINした結果、各カラム同士を比較することとなる。 SELECT ... FROM input_tables as l CROSS JOIN input_tables as r あるカラムが条件に合わなければ、もうその先は見ても意味がない、というケースは多い。例えば、まず first_name 、surname が同じでなければ、その先の比較を行わない、というのはあり得る。 SELECT ... FROM input_tables as l INNER JOIN input_tables as r ON l.first_name = r.first_name AND l.surname = r.surname このような考え方をブロッキング、ON句の条件をブロッキングルールと言う。ただ、これだと性と名が完全一致していないレコードが残らない。そこで、ブロッキングルールを複数定義し、いずれかが真であれば残すことができる。ここでポイントなのが、ブロッキングルールを複数定義したとき、それぞれのブロッキングルールで重複して選ばれるレコードが発生した場合、 Splinkが自動的に排除してくれる。このため、ブロッキングルールを重ねがけすると、最終的に残るレコード数は一致する。ただ、順番により、同じルールで残るレコード数は変化する。逆に言うと、ブロッキングルールを足すことで、重複除去後のOR条件が増えていく。積算グラフにして、ブロッキングルールとその順番の効果を見ることができる。 from splink import DuckDBAPI, block_on from splink.blocking_analysis import ( cumulative_comparisons_to_be_scored_from_blocking_rules_chart, ) blocking_rules = [ block_on(\"substr(first_name,1,3)\", \"substr(surname,1,4)\"), block_on(\"surname\", \"dob\"), block_on(\"first_name\", \"dob\"), block_on(\"postcode_fake\", \"first_name\"), block_on(\"postcode_fake\", \"surname\"), block_on(\"dob\", \"birth_place\"), block_on(\"substr(postcode_fake,1,3)\", \"dob\"), block_on(\"substr(postcode_fake,1,3)\", \"first_name\"), block_on(\"substr(postcode_fake,1,3)\", \"surname\"), block_on(\"substr(first_name,1,2)\", \"substr(surname,1,2)\", \"substr(dob,1,4)\"), ] db_api = DuckDBAPI() cumulative_comparisons_to_be_scored_from_blocking_rules_chart( table_or_tables=df, blocking_rules=blocking_rules, db_api=db_api, link_type=\"dedupe_only\", ) 積算グラフは以下の通り。積み上がっている数値は「比較の数」。要は、論理和で条件を足していって、次第に緩和されている様子がわかる。 DuckDBでは比較の数を2,000万件以内、Athena,Sparkでは1億件以内を目安にせよとのこと。比較の定義 Splinkは Fellegi-Sunter model モデル (というかフレームワーク) に基づいている。 https://moj-analytical-services.github.io/splink/topic_guides/theory/fellegi_sunter.html 各カラムの同士をカラムの特性に応じた距離を使って比較し、重みを計算していく。各カラムの比較に使うためのメソッドが予め用意されているので、特性に応じて選んでいく。以下では、first_name, sur_name に ForenameSurnameComparison が使われている。 dobにDateOfBirthComparison、birth_place、ocupationにExactMatchが使われている。 import splink.comparison_library as cl from splink import Linker, SettingsCreator settings = SettingsCreator( link_type=\"dedupe_only\", blocking_rules_to_generate_predictions=blocking_rules, comparisons=[ cl.ForenameSurnameComparison( \"first_name\", \"surname\", forename_surname_concat_col_name=\"first_name_surname_concat\", ), cl.DateOfBirthComparison( \"dob\", input_is_string=True ), cl.PostcodeComparison(\"postcode_fake\"), cl.ExactMatch(\"birth_place\").configure(term_frequency_adjustments=True), cl.ExactMatch(\"occupation\").configure(term_frequency_adjustments=True), ], retain_intermediate_calculation_columns=True, ) # Needed to apply term frequencies to first+surname comparison df[\"first_name_surname_concat\"] = df[\"first_name\"] + \" \" + df[\"surname\"] linker = Linker(df, settings, db_api=db_api) ComparisonとComparison Level ここでSplinkツール内の比較の概念の説明。以下の通り概念に名前がついている。 Data Linking Model ├─-- Comparison: Date of birth │ ├─-- ComparisonLevel: Exact match │ ├─-- ComparisonLevel: One character difference │ ├─-- ComparisonLevel: All other ├─-- Comparison: First name │ ├─-- ComparisonLevel: Exact match on first_name │ ├─-- ComparisonLevel: first_names have JaroWinklerSimilarity > 0.95 │ ├─-- ComparisonLevel: first_names have JaroWinklerSimilarity > 0.8 │ ├─-- ComparisonLevel: All other モデルのパラメタ推定モデルの実行に必要なパラメタは以下の3つ。Splinkを用いてパラメタを得る。ちなみに u は \"\'U\'nmatch\"、m は \"\'M\'atch\"。背後の数式の説明で現れる。 No パラメタ説明 1 無作為に選んだレコードが一致する確率入力データからランダムに取得した2つのレコードが一致する確率 (通常は非常に小さい数値) 2 u値(u確率) 実際には一致しないレコードの中で各 ComparisonLevel に該当するレコードの割合。具体的には、レコード同士が同じエンティティを表すにも関わらず値が異なる確率。例えば、同じ人なのにレコードによって生年月日が違う確率。これは端的には「データ品質」を表す。名前であればタイプミス、別名、ニックネーム、ミドルネーム、結婚後の姓など。 3 m値(m確率) 実際に一致するレコードの中で各 ComparisonLevel に該当するレコードの割合。具体的には、レコード同士が異なるエンティティを表すにも関わらず値が同じである確率。例えば別人なのにレコードによって性・名が同じ確率 (同姓同名)。性別は男か女かしかないので別人でも50%の確率で一致してしまう。無作為に選んだレコードが一致する確率入力データからランダムに抽出した2つのレコードが一致する確率を求める。値は0.000136。すべての可能なレコードのペア比較のうち7,362.31組に1組が一致すると予想される。合計1,279,041,753組の比較が可能なため、一致するペアは合計で約173,728.33組になると予想される、とのこと。 linker.training.estimate_probability_two_random_records_match( [ block_on(\"first_name\", \"surname\", \"dob\"), block_on(\"substr(first_name,1,2)\", \"surname\", \"substr(postcode_fake,1,2)\"), block_on(\"dob\", \"postcode_fake\"), ], recall=0.6, ) > Probability two random records match is estimated to be 0.000136. > This means that amongst all possible pairwise record comparisons, > one in 7,362.31 are expected to match. > With 1,279,041,753 total possible comparisons, > we expect a total of around 173,728.33 matching pairs u確率の推定実際には一致しないレコードの中でComparisonの評価結果がPositiveである確率。基本、無作為に抽出したレコードは一致しないため、「無作為に抽出したレコード」を「実際には一致しないレコード」として扱える、という点がミソ。 probability_two_random_records_match によって得られた値を使ってu確率を求める。 estimate_u_using_random_sampling によって、ラベルなし、つまり教師なしでu確率を得られる。レコードのペアをランダムでサンプルして上で定義したComparisonを評価する。ランダムサンプルなので大量の不一致が発生するが、各Comparisonにおける不一致の分布を得ている。これは、例えば性別について、50%が一致、50%が不一致である、という分布を得ている。一方、例えば生年月日について、一致する確率は 1%、1 文字の違いがある確率は 3%、その他はすべて 96% の確率で発生する、という分布を得ている。 linker.training.estimate_u_using_random_sampling(max_pairs=5e6) > ----- Estimating u probabilities using random sampling ----- > > Estimated u probabilities using random sampling > > Your model is not yet fully trained. Missing estimates for: > - first_name_surname (no m values are trained). > - dob (no m values are trained). > - postcode_fake (no m values are trained). > - birth_place (no m values are trained). > - occupation (no m values are trained). m確率の推定「実際に一致するレコード」の中で、Comparisonの評価がNegativeになる確率。そもそも、このモデルを使って名寄せ、つまり「一致するレコード」を見つけたいのだから、モデルを作るために「実際に一致するレコード」を計算しなければならないのは矛盾では..となる。無作為抽出結果から求められるu確率とは異なり、m確率を求めるのは難しい。もしラベル付けされた「一致するレコード」、つまり教師データセットがあるのであれば、そのデータセットを使ってm確率を求められる。例えば、日本人全員にマイナンバーが振られて、全てのレコードにマイナンバーが振られている、というアナザーワールドがあるのであれば、マイナンバーを使ってm確率を推定する。(どういう状況??) ラベル付けされたデータがないのであれば、EMアルゴリズムでm確率を求めることになっている。 EMアルゴリズムは反復的な手法で、メモリや収束速度の点でペア数を減らす必要があり、例ではブロッキングルールを設定している。以下のケースでは、first_nameとsurnameをブロッキングルールとしている。つまり、first_name, surnameが完全に一致するレコードについてペア比較を行う。この仮定を設定したため、first_name, surname (first_name_surname) のパラメタを推定できない。 training_blocking_rule = block_on(\"first_name\", \"surname\") training_session_names = ( linker.training.estimate_parameters_using_expectation_maximisation( training_blocking_rule, estimate_without_term_frequencies=True ) ) > ----- Starting EM training session ----- > > Estimating the m probabilities of the model by blocking on: > (l.\"first_name\" = r.\"first_name\") AND (l.\"surname\" = r.\"surname\") > > Parameter estimates will be made for the following comparison(s): > - dob > - postcode_fake > - birth_place > - occupation > > Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: > - first_name_surname > > Iteration 1: Largest change in params was 0.248 in probability_two_random_records_match > Iteration 2: Largest change in params was 0.0929 in probability_two_random_records_match > Iteration 3: Largest change in params was -0.0237 in the m_probability of birth_place, level `Exact match on > birth_place` > Iteration 4: Largest change in params was 0.00961 in the m_probability of birth_place, level `All other >comparisons` > Iteration 5: Largest change in params was -0.00457 in the m_probability of birth_place, level `Exact match on birth_place` > Iteration 6: Largest change in params was -0.00256 in the m_probability of birth_place, level `Exact match on birth_place` > Iteration 7: Largest change in params was 0.00171 in the m_probability of dob, level `Abs date difference Iteration 8: Largest change in params was 0.00115 in the m_probability of dob, level `Abs date difference Iteration 9: Largest change in params was 0.000759 in the m_probability of dob, level `Abs date difference Iteration 10: Largest change in params was 0.000498 in the m_probability of dob, level `Abs date difference Iteration 11: Largest change in params was 0.000326 in the m_probability of dob, level `Abs date difference Iteration 12: Largest change in params was 0.000213 in the m_probability of dob, level `Abs date difference Iteration 13: Largest change in params was 0.000139 in the m_probability of dob, level `Abs date difference Iteration 14: Largest change in params was 9.04e-05 in the m_probability of dob, level `Abs date difference <= 10 year` 同様にdobをブロッキングルールに設定して実行すると、dob以外の列についてパラメタを推定できる。 training_blocking_rule = block_on(\"dob\") training_session_dob = ( linker.training.estimate_parameters_using_expectation_maximisation( training_blocking_rule, estimate_without_term_frequencies=True ) ) > ----- Starting EM training session ----- > > Estimating the m probabilities of the model by blocking on: > l.\"dob\" = r.\"dob\" > > Parameter estimates will be made for the following comparison(s): > - first_name_surname > - postcode_fake > - birth_place > - occupation > > Parameter estimates cannot be made for the following comparison(s) since they are used in the blocking rules: > - dob > > Iteration 1: Largest change in params was -0.474 in the m_probability of first_name_surname, level `Exact match on first_name_surname_concat` > Iteration 2: Largest change in params was 0.052 in the m_probability of first_name_surname, level `All other comparisons` > Iteration 3: Largest change in params was 0.0174 in the m_probability of first_name_surname, level `All other comparisons` > Iteration 4: Largest change in params was 0.00532 in the m_probability of first_name_surname, level `All other comparisons` > Iteration 5: Largest change in params was 0.00165 in the m_probability of first_name_surname, level `All other comparisons` > Iteration 6: Largest change in params was 0.00052 in the m_probability of first_name_surname, level `All other comparisons` > Iteration 7: Largest change in params was 0.000165 in the m_probability of first_name_surname, level `All other comparisons` > Iteration 8: Largest change in params was 5.29e-05 in the m_probability of first_name_surname, level `All other comparisons` > > EM converged after 8 iterations > > Your model is not yet fully trained. Missing estimates for: > - first_name_surname (some u values are not trained). モデルパラメタの可視化 m確率、u確率の可視化。マッチウェイトの可視化。マッチウェイトは (log_2 (m / u))で計算される。 linker.visualisations.match_weights_chart() モデルの保存と読み込み以下でモデルを保存できる。 settings = linker.misc.save_model_to_json( \"./saved_model_from_demo.json\", overwrite=True ) 以下で保存したモデルを読み込める。 import json settings = json.load( open(\'./saved_model_from_demo.json\', \'r\') ) リンクするのに十分な情報が含まれていないレコード「John Smith」のみを含み、他のすべてのフィールドがnullであるレコードは、他のレコードにリンクされている可能性もあるが、潜在的なリンクを明確にするには十分な情報がない。以下により可視化できる。 linker.evaluation.unlinkables_chart() 横軸は「マッチウェイトの閾値」。縦軸は「リンクするのに十分な情報が含まれないレコード」の割合。マッチウェイト閾値=6.11ぐらいのところを見ると、入力データセットのレコードの約1.3%がリンクできないことが示唆される。訓練済みモデルを使って未知データのマッチウェイトを予測上で構築した推定モデルを使用し、どのペア比較が一致するかを予測する。内部的には以下を行うとのこと。 blocking_rules_to_generate_predictionsの少なくとも1つと一致するペア比較を生成 Comparisonで指定されたルールを使用して、入力データの類似性を評価推定された一致重みを使用し、要求に応じて用語頻度調整を適用して、最終的な一致重みと一致確率スコアを生成 df_predictions = linker.inference.predict(threshold_match_probability=0.2) df_predictions.as_pandas_dataframe(limit=1) > Blocking time: 0.88 seconds > Predict time: 1.91 seconds > > -- WARNING -- > You have called predict(), but there are some parameter estimates which have neither been estimated or > specified in your settings dictionary. To produce predictions the following untrained trained parameters will > use default values. > Comparison: \'first_name_surname\': > u values not fully trained records_to_plot = df_e.to_dict(orient=\"records\") linker.visualisations.waterfall_chart(records_to_plot, filter_nulls=False) predictしたマッチウェイトの可視化、数式との照合 predictしたマッチウェイトは、ウォーターフォール図で可視化できる。マッチウェイトは、モデル内の各特徴量によって一致の証拠がどの程度提供されるかを示す中心的な指標。 (lambda)は無作為抽出した2つのレコードが一致する確率。(K=m/u)はベイズ因子。 begin{align} M &= log_2 ( frac{lambda}{1-lambda} ) + log_2 K \\ &= log_2 ( frac{lambda}{1-lambda} ) + log_2 m - log_2 u end{align} 異なる列の比較が互いに独立しているという仮定を置いていて、 2つのレコードのベイズ係数が各列比較のベイズ係数の積として扱う。 begin{eqnarray} K_{feature} = K_{first_name_surname} + K_{dob} + K_{postcode_fake} + K_{birth_place} + K_{occupation} + cdots end{eqnarray} マッチウェイトは以下の和。 begin{eqnarray} M_{observe} = M_{prior} + M_{feature} end{eqnarray} ここで begin{align} M_{prior} &= log_2 (frac{lambda}{1-lambda}) \\ M_{feature} &= M_{first_name_surname} + M_{dob} + M_{postcode_fake} + M_{birth_place} + M_{occupation} + cdots end{align} 以下のように書き換える。 begin{align} M_{observe} &= log_2 (frac{lambda}{1-lambda}) + sum_i^{feature} log_2 (frac{m_i}{u_i}) \\ &= log_2 (frac{lambda}{1-lambda}) + log_2 (prod_i^{feature} (frac{m_i}{u_i}) ) end{align} ウォーターフォール図の一番左、赤いバーは(M_{prior} = log_2 (frac{lambda}{1-lambda}))。特徴に関する追加の知識が考慮されていない場合のマッチウェイト。横に並んでいる薄い緑のバーは (M_{first_name_surname} + M_{dob} + M_{postcode_fake} + M_{birth_place} + M_{occupation} + cdots)。各特徴量のマッチウェイト。一番右の濃い緑のバーは2つのレコードの合計マッチウェイト。 begin{align} M_{feature} &= M_{first_name_surname} + M_{dob} + M_{postcode_fake} + M_{birth_place} + M_{occupation} + cdots \\ &= 8.50w end{align} まとめ長くなったのでいったん終了。この記事では教師なし確率的名寄せパッケージSplinkを使用してモデルを作ってみた。次の記事では、作ったモデルを使用して実際に名寄せをしてみる。途中、DuckDBが楽しいことに気づいたので、DuckDBだけで何個か記事にしてみようと思う。

Terraformを使ってAWSにWebアプリケーションの実行環境を立てる (EC2立てるまで)

この記事で紹介する範囲

Terraformの導入

git secretsの導入

ディレクトリ構成

tfstateの保存先の定義

credentialsの書き方

providerの定義

エントリポイント

VPCモジュール

EC2モジュール

実行

疎通確認

React+Next.jsでDummy JSONのCRUDをCSR/SSRの両方で作成して違いを調べてみた話

go-txdbを使ってgolang, gin, gorm(gen)+sqlite構成のAPI をテストケース毎に管理する

gorm互換の型安全なORMであるgenでCRUD APIを試作

Golang + Gin カスタムバリデーション

Golang + Gin Framework で Hello World してみた話〜基本的なルーティング、バスパラメタ・クエリパラメタ・JSON Req/Res、フォームデータ

Snowflake MCPサーバを試してみた

Fellegi-Sunterモデルに基づく確率的名寄せパッケージ Splinkを試してみる

AirflowでEnd-To-End Pipeline Testsを行うためにAirflow APIを調べてみた話

CustomOperatorのUnitTestを理解するためGCSToBigQueryOperatorのUnitTestを読んでみた話

GoogleによるAirflow DAG実装のベスプラ集を読んでみた – その1

この記事で紹介する範囲

Terraformの導入

git secretsの導入

ディレクトリ構成

tfstateの保存先の定義

credentialsの書き方

providerの定義

エントリポイント

VPCモジュール

EC2モジュール

実行

疎通確認

関連記事