You are coding as a Guest. Sign in with your RoleNest account to permanently track your streak, earn XP, and climb the Campus Leaderboard!
Sign In with RoleNestProblem Set
🔥BPE Token Pair Frequency (Tokenizer Engine)Medium
MediumAI & Machine Learning•Acceptance: 61.3%
BPE Token Pair Frequency (Tokenizer Engine)
Real-World Engineering Context
Core subword vocabulary generation step used by OpenAI's tiktoken and Hugging Face tokenizers.
In Byte-Pair Encoding (BPE), tokenizers repeatedly merge the most frequent adjacent character pair. Given an array of space-separated token strings `words` and a target 2-element pair `[char1, char2]`, count how many times this exact adjacent pair appears across all words.
Sample Test Cases
Input: [["l o w","l o w e r"],["l","o"]]
Expected: 2
Input: [["a b a b","a b c"],["a","b"]]
Expected: 3
Constraints
- 1 <= words.length <= 1000
- pair.length == 2
Language:
Ready to test. Click Run Code or Submit Solution to run test cases in isolated browser sandbox.