Hey guys, what is a good number of tablets in a bundle? is too many tablets bad?
I have this setup with miranda and mercuries, and using this in-house streaming platform to write into dynamic tables (a write heavy workload), but constantly get alerts with delays, AI says that it could be from too many cells per tablet instance and too many tablets, is this true?
For example on mercuries i have 7k tablets, across 120 cells, 10 cells per instance on mercury, each instance is like 16 cpu cores and 64 gigs of ram. I store like less than 1TB of data so far, maybe i'll store like 10tb or so in the near future.
How many tablets should i have and how many cells per instance?
Should i also increase dynamic memory and keep static memory low if i have a write heavy workload?
И Hi! Our best practice recommendation is to have tablets of about 5-10 GB compressed size and to have 5-10 tablet cells per node. Exact numbers do not matter too much, though there may be problems at the extreme points: too large tablets (100+ GB) make LSM algorithms less efficient, and too small tablets (< 100 MB) may induce overhead because of small chunks and dynamic stores. Other than that, the setup is rather flexible. Your setup looks pretty much OK. Regarding the memory — I would not recommend increasing tablet dynamic that much because it increases recovery time and does not help write throughput at all.
Photo
click to show
click to show
Okay, but for some reason I get network delay alerts, they look something like this
On this particular instance i am showing you, there's nothing above quota thresholds in regards to CPU/RAM/Network usage, the only graphs out of the ordinary are the ones called "tablet node write data weight rate"
Photo
click to show
click to show
Hello guys. Could you please explain this to me?
it says that the first min_data_versions values cannot be deleted, but then it says "at least one value (the last one) is always stored"
i don't get it, this doc looks confusing
if i have:
min_data_versions=4
max_data_versions=5
how many will i end up with after compaction, 5 or 4?
А A $test = [1, 2, 3, 1, 4, 1, 3]; $dictCount = ($list) -> { $tuples = ListMap($list, ($item)->(AsTuple($item, Void()))); $tuplesWithCount = ListMap(DictItems(ToMultiDict($tuples)), ($tuple) -> (AsTuple($tuple.0, ListLength($tuple.1)))); return ToDict($tuplesWithCount); }; select $dictCount($test); More efficient method involves mutable dicts which are not available yet in released Query Tracker
$data = [
<|id: 1, val: "a"|>,
<|id: 1, val: "b"|>,
<|id: 1, val: "a"|>,
<|id: 2, val: "c"|>,
<|id: 2, val: "a"|>,
<|id: 2, val: "b"|>,
];
SELECT DictAggregate(
ToMultiDict(
ListNotNull(AGGREGATE_LIST(
IF(val IS NOT NULL, AsTuple(val, 1), NULL)
))
),
AggregationFactory("COUNT")
) AS counts
from AS_TABLE($data)
group by id;
this seemed to work
alright thanks
I used this (modified, from an earlier answer):
$GetItemCounts = ($item_list) -> {
$initial_dict = CAST({} as Dict<String, Int32>);
RETURN ListFold(
$item_list,
$initial_dict,
($current_item, $acc_counts_dict) -> {
$single_item_dict = {$current_item: 1};
RETURN SetUnion(
$acc_counts_dict,
$single_item_dict,
($key, $existing_count, $new_count) -> (
($existing_count ?? 0) + ($new_count ?? 0)
)
);
}
);
};
И M