Paper deep dive
Weight and Height Estimation from a Single Human Image Captured in the Wild
Hira Yaseen, Arif Mahmood, Waqas Sultani
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/3/2026, 10:45:50 AM
Summary
This paper addresses the challenge of estimating Body Mass Index (BMI), weight, and height from single human images captured in the wild. The authors propose a new dataset, ITU-BMI, containing 6,105 images with ground truth labels, collected from social networking sites. They employ deep neural networks (VGG, DenseNet, ResNet) using single and multi-task learning approaches with various input modalities (RGB, depth, pose-affinity, edge maps). Experimental results indicate that full-body images yield better estimation accuracy than half-body or facial images, and multi-task learning generally outperforms single-task learning.
Entities (11)
Relation Signals (11)
ITU-BMI → contains → 6105 images
confidence 98% · we propose a new dataset consisting of 6105 images with ground truth labels
ITU-BMI → provideslabelsfor → Height
confidence 95% · ground truth labels of height, weight and BMI
ITU-BMI → provideslabelsfor → BMI
confidence 95% · ground truth labels of height, weight and BMI
ITU-BMI → provideslabelsfor → Weight
confidence 95% · ground truth labels of height, weight and BMI
Full Body Images → producesbetterresultsthan → Half Body Images
confidence 92% · full body images have produced better results than the other half body and facial images
Full Body Images → producesbetterresultsthan → Facial Images
confidence 92% · full body images have produced better results than the other half body and facial images
Masked RCNN → usedfor → Human Separation
confidence 90% · We use Masked RCNN [37] to separate the humans
VGG → usedfor → BMI Estimation
confidence 90% · extensive experimentation is performed using different CNN backbones including VGG
→ →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:A person's physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and finances. Body Mass Index (BMI) is a well known measure that encodes the characteristics of both the weight and the height. BMI has been used as a self-monitoring tool, and it has long-term implications on one's life. For example, it may help predicting the risk of various diseases and estimating longevity. Automatic BMI estimation using a single person image in the wild is a challenging task due to wide variations in human pose, camera geometry, personal appearance and distracting backgrounds. In this paper, we explore the performance of deep neural networks using single and multi-task learning by employing different modalities including RGB, depth-maps, pose-affinity maps, and edge-maps to predict BMI, weight, and height from daily life images available on social networking websites. Currently, no full body image dataset for BMI estimation is publicly available, therefore we propose a new dataset consisting of 6105 images with ground truth labels of height, weight and BMI. Our proposed dataset is collected in the wild containing images from various ethnicity and distributed over varying age groups and gender. It consists of frontal, back, full and half body, side poses, mirror selfies with varying backgrounds and scale variations and may contain artifacts hiding partial or full face. Extensive experimentation is performed using full body, half body and face images only using different CNN backbones including VGG, Densenet and ResNet. Our experimental results demonstrate that full body images have produced better results than the other half body and facial images in the wild.
Tags
Links
- Source: https://arxiv.org/abs/2607.26104v1
- Canonical: https://arxiv.org/abs/2607.26104v1
Trouble viewing inline? Open PDF directly →
Full Text
68,255 characters extracted from source content.
Expand or collapse full text
Weight and Height Estimation from a Single Human Image Captured in the Wild Hira Yaseen, Arif Mahmood Waqas Sultani H. Yaseen, A. Mahmood, W. Sulatni are with the Department of Computer Science, Information Technology University (ITU), 346-B, Ferozepur Road, Lahore, Pakistan. E-mails: PhDCS17002@itu.edu.pk, arif.mahmood@itu.edu.pk, waqas.sultani@itu.edu.pk Abstract A person’s physical characteristics such as weight and height are important indicators of his physical and mental health, daily life routines and finances. Body Mass Index (BMI) is a well known measure that encodes the characteristics of both the weight and the height. BMI has been used as a self-monitoring tool, and it has long-term implications on one’s life. For example, it may help predicting the risk of various diseases and estimating longevity. Automatic BMI estimation using a single person image in the wild is a challenging task due to wide variations in human pose, camera geometry, personal appearance and distracting backgrounds. In this paper, we explore the performance of deep neural networks using single and multi-task learning by employing different modalities including RGB, depth-maps, pose-affinity maps, and edge-maps to predict BMI, weight, and height from daily life images available on social networking websites. Currently, no full body image dataset for BMI estimation is publicly available, therefore we propose a new dataset consisting of 6105 images with ground truth labels of height, weight and BMI. Our proposed dataset is collected in the wild containing images from various ethnicity and distributed over varying age groups and gender. It consists of frontal, back, full and half body, side poses, mirror selfies with varying backgrounds and scale variations and may contain artifacts hiding partial or full face. Extensive experimentation is performed using full body, half body and face images only using different CNN backbones including VGG, Densenet and ResNet. Our experimental results demonstrate that full body images have produced better results than the other half body and facial images in the wild. Also in most of the experiments, joint learning of height, weight and BMI through multi-task learning has performed better than single task learning. The code and the dataset will soon be made publicly available. I Introduction Analyzing physical characteristics of a person such as weight and height are of much interest to doctors and researchers as well as to the general public in the current internet age. In order to capture various combinations of height and weight, Body Mass Index (BMI) [40, 39] has been introduced as a useful measure which is defined as BMI = weight/height2, when weight is measured in kilograms and height in meters. Healthcare professionals have divided BMI into four broad categories: underweight, normal, overweight, and obese. If BMI <18.5<18.5, he/she is underweight, if 18.5 ≤ BMI <25<25 then he/she is normal, if 25≤25≤ BMI <30<30 then he/she is overweight, and BMI ≥30≥ 30 is considered as obese. A healthy person must maintain a normal BMI, because an increased BMI is correlated with many diseases such as type-2 diabetes [9, 10, 55], gallbladder diseases, asthma [12], sleep apnea [11], high blood pressure, angina [13], fatty liver, cardiovascular diseases [56], osteoarthritis, heart strokes and cancer [8, 57]. Similarly, an underweight person may suffer from anorexia, type-1 diabetes, hyperthyroidism, amenorrhea, infertility, osteoporosis [14], immunodeficiency [16, 15], and tuberculosis. Instead of physically measuring a persons height and weight, various techniques have been employed by medical practitioners to indirectly estimate a person’s body composition and his BMI. These methods include Bioelectrical Impedance Analysis (BIA) [61, 62], Bioimpedance Spectroscopy, Dual-energy X-ray absorptiometry (DXA) [63, 72], Air Displacement Plethysmography (ADP) [64, 65], three-dimensional Photonic Scanner (3D-PS) [66, 67], Magnetic Resonance Spectroscopy (MRS) [68], and Positron Emission Tomography (PET) [69]. In contrast to these techniques, our proposed method provides an easy and efficient way to automatically estimate BMI of a person using his/her daily life pictures captured in the wild. On many social networking sites 111Facebook, twitter, Instagram, , people like to share pictures of their important life events. On some health and weight loss/gain related sites 222https://w.reddit.com/r/progresspics/, https://myprogresspics.com/, people share their pictures as well as corresponding weights and heights to become an inspiration for others as well as for themselves. In this work we utilize such labelled pictures for training deep neural networks to predict a person’s weight, height and BMI from his/her pictures captured without a controlled environment. Once our networks are trained then we can employ the trained networks to predict these attributes of persons on social platforms at a mass scale. Our algorithm may potentially be used to monitor public health on large geographical areas [7, 26]. Such as system can also be used for self health assessments as well as to warn family and friends about their health conditions. Furthermore, estimating height, and weight of a person can have significant applications in person re-identification, forensics, crime scene investigations, and surveillance and security systems. Due to the broad scope of BMI estimation from images, many researchers have proposed artificial intelligent and machine learning based approach for this purpose [6, 5, 4]. However, most of these methods require images to be captured under control environment limiting their application to relatively narrow scope. In contrast, we propose a method to estimate BMI using images captured in the wild from different viewing conditions, camera resolutions, various scales, as well as selfies. Many existing methods estimate BMI using frontal face images captured under controlled environment. These images exhibit small variation and have a very constrained distribution of data. To address this issue, Jiang et al., [6] collected paired frontal full-body images of the same person having different BMI. They extracted hand-crafted features to train support vector regressor (SVR) and Gaussian process regressor (GPR) to predict BMI. Although the authors reported good results, however, they have not made their dataset publicly available limiting its application. In contrast, we propose a dataset having a large number of human body images with the no-pair condition captured in the wild. Our collected dataset has significant variations and contains frontal/side poses, faces/full body/upper body images. We will soon make this dataset publicly available. We train deep neural networks including VGG-16, DensNet-121 and ResNet-50 in multi-task fashion as well as single task learning for the estimation of height, weight, and BMI. We also pose the problem as classification of under-weight, normal, over-weight and obese. In addition to RGB and gray scale image representations, we have thoroughly experimented with pose affinity maps, depth maps and human segmentation to improve BMI estimation. Our experimental results have demonstrated significantly improved performance as compared to the existing state of the art methods. The main contributions of this work include: • We propose a large scale publicly available data set of labeled 6105 images • We propose the use of multiple modalities such as pose affinity maps, depth maps and human segmentation, in addition to RGB and gray scale image representations for BMI estimation. • We propose joint learning of weight, height and BMI through multi-task learning using deep neural networks. The rest of the paper is organized as follows: Section 2 gives a brief overview of the existing methods to estimate weight, height, and/or BMI from faces or body images. Section 3 covers our proposed methodology including multi-task learning, single task learning and different modalities to improve weight, height and BMI estimations. Section 4 contains experiments and results while Section 5 concludes the paper. I Review of Existing BMI Estimation Methods Due to a large scale applications of BMI for the prediction of health conditions, several automatic and efficient methods have been proposed. These methods include BMI estimation from facial and full body images using hand crafted features and deep neural networks. We broadly categorize these methods as follows: I-A BMI Estimation Using Facial Images Due to the availability of face images in social media websites such as Facebook, Instagram and Twitter, recently, there has been a growing interest in predicting a person’s BMI from his face image. I-A1 Face to BMI estimation using hand-crafted features Wen et al. [18] used hand-crafted features to estimate BMI of a person using least squares regression, support vector regression (SVR), and gaussian process regression (GPR) from only frontal facial images in MORPH-I dataset [30]. The authors used seven hand-crafted facial features including cheekbone to jaw width ratio, width to upper facial height ratio, perimeter to area ratio, eye size, lower face to face height ratio, face width to lower face height ratio and mean of eyebrow height, which were detected automatically using active shape model. They have reported good results, however, their dataset/code is not publicly available. Barr et al. [27] extracted hand-crafted features using facial landmarks and employed SVR for BMI estimations. Their dataset was captured in a controlled environment and consisted of frontal facial images in neutral emotion. They used the same seven facial features as proposed by [18]. Andersson et al. [47] used anthropometric features and gait information to classify gender and BMI estimation on 106 persons whose images were taken using Microsoft Kinect sensor. Different combinations of attributes from the extracted anthropometric features were used and several classifiers such as SVM, K-nearest neighbour (KNN), and multi-layer perceptron (MLP) were employed. Pascali et al. [50] proposed a method to automatically extract geometric features from 3D facial images (using depth scanners) and predicted body weight. Affuso et al. [46] collected data from a selected set of 343 participants under controlled environment. For each participant, three full body images were captured including frontal, back and side poses. They extracted four features from the captured images including body volume, front curve, side curve, and body shape which are then fed to SVR to predict body fatness. Farina et al. [48] introduced the use of smartphones camera images to estimate body fat mass. They collected a data set (digital images, and DXA) of 117 subjects. Estimating body fat mass can help in weight management and BMI estimation. I-A2 Face to BMI estimation using Deep-Nets Kocabey et al. [5] used VGG-Face and VGG-Net to extract facial features and epsilon SVR for BMI estimation over a dataset of facial images captured in the wild. They achieved excellent results compared to hand-crafted features. In their dataset, Americans and Africans were mostly obese, therefore their algorithm training is considered biased due to high prior probability. Jahandideh et al. [32] trained ResNet-50 on the VGG-Face dataset and then fine-tuned on Faces in the Real World (FIRW) dataset to estimate body type, ethnicity, gender, height, and weight. A separate ResNet-50 was trained for height and weight classification. Jiang et al. [33] predicted BMI from face images using deep features with SVR. They compared results of VGG-Face, Light CNN, center loss [35], and Arcface [36]. Dantcheva et al. [4] used ResNet-50 to predict height, weight, and BMI from the VIP attribute face images dataset. Three separate networks were trained for height, weight, and BMI estimation. One limitation of their work is clear frontal face images of 1026 celebrities cropped using the Viola-Jones algorithm [29]. Another limitation is a biased dataset that contains 22 ≤ BMI ≤ 28 for around 80% samples. Similarly, Jiang et al. [51] estimated BMI from facial images using a label distribution. BMI was estimated as a discrete probability distribution and estimation results were demonstrated on FIW-BMI [33], Morph-I, and VIP-attribute data sets. I-B BMI Estimation using Full Body Images Some of the recent studies have demonstrated the estimation of BMI from 2D and 3D images. Below, we discuss them briefly. I-B1 Body to BMI estimation using Hand-crafted Features Jiang and Guo [6] work computes five anthropometric features (waist width to thigh width, waist width to hip width, waist width to head width, hip-width to head width, and the area between waist and hip) obtained from Body contour and skeleton joints (CSJ) detection network on frontal full-body images. After extracting anthropometric features, the authors apply support vector regression (SVR) and gaussian process regression (GPR) methods to predict BMI directly from images. Velardo et al. [19] have collected anthropometric data from the National Health and Nutrition Examination Survey (NHANES) to estimate human body weight. However, the authors did not consider samples that exhibit weights below 35kg. and beyond 130 kilograms. The anthropometric features obtained from [24] are fed directly to regression models for weight estimation. The anthropometric features used by this method include height, upper leg length, calf circumference, upper arm length, upper arm circumference, waist circumference, and upper leg circumference. The authors also collected another dataset of 40 pictures (a frontal shot and a profile one) taken at a fixed distance from the camera and provided results as real-case analysis. Pfitzner et al. [25] provides weight estimation on a dataset of 233 people. The authors computed anthropometric features from the frontal body RGB-D sensor data, and then employed artificial neural network (ANN) for weight estimation. Temperature features from a thermal camera are also included to better estimate the weight of a human body. In the subsequent work [45], Pfitzner et al. proposed an approach to predict the weight of a person using RGBD camera sensor photos in three different poses: lying, walking, and standing. I-B2 Body to BMI estimation using Deep-Nets Nahavandi et al. [23] proposed to use ResNet-18 & ResNet-34 to extract deep features and present a skeleton-free Kinect system to estimate the human body mass index. Specifically, the authors generated depth images using a Kinect camera and encoded them into RGB for ResNet fine-tuning. The generated depth images are translated into a 3D point cloud to calculate the body surface area (BSA). BSA and other anthropometric features are then used to estimate the weight. Vakli et al. [52], for the first time, investigated structural brain images/MRI to predict body mass index of a person. The authors used a CNN on MRI along with age and sex information to estimate BMI. Gunel et al. [49] have collected a new 2D images data set with unknown camera parameters, and scene geometry. The authors estimate the height of a person using pose, bone-lengths ratios, and facial features using different linear, shallow, and deep machine learning algorithms. In contrast to the above mentioned works, we propose to publicly release a large scale full body dataset for BMI estimation. To the best of our knowledge, no such dataset is publicly available yet. We performed extensive experiments on this dataset and demonstrated that, in general, joint estimation of height, weight and BMI performs better than that of their estimation in the separate frameworks. Our thorough analysis using different modalities including RGB, depth-maps, pose-affinity maps, and masks employing various deep networks validate the proposed ideas. I Review of Existing Image to BMI estimation Datasets In this section, we briefly review image to BMI estimation datasets which are discussed in the literature. Table I shows a comparison of our proposed Body2BMI-ITU dataset with the other datasets which were proposed for the similar problem. I-A Face2BMI Dataset Kocabey et al. [5] proposed a publicly available Face2BMI dataset for the estimation of BMI from face images. This dataset contains annotated faces of 2103 pairs from reddit.com, with a total number of 4206 images with past and current weights. This dataset contains 2438 males and 1768 females. The authors define seven BMI categories where 7 subjects have underweight BMI, 680 have normal BMI, 1151 have overweight BMI, 941 are moderately obese, 681 are severely obese, and 746 are very severely obese. I-B FIW-BMI Dataset Jiang et al. [33] proposed this dataset which contains 7930 wild face images of 4881 individuals along with the labels of gender, height, and weight. In this dataset, there exist 3192 males (5197 images) and 1689 females (2733 images). Among these face images, 43 are underweight, 1662 are normal, 2455 are overweight, and 3770 are obese. This dataset is not publicly available. I-C Morph-I Data set The Morph I database [30] contains 55608 mugshot-style frontal view face images of 13673 subjects of the age between 16 and 99 years where 47057 images correspond to male persons and 8551 to female persons. The dataset has the age, gender, ethnicity, weight and height labels and 55% of the data lies under ’normal’ BMI range. I-D VIP-Attribute Dataset This publicly available dataset contains 1026 celebrities’ frontal face images. It contain 523 male and equal number of female high resolution cropped images with no background. TABLE I: A comparison of existing datasets with proposed datasets: ITU-BMI, ITU-BMI_F (only face images), ITU-BMI_U (all images with upper-body), ITU-BMI_B (only images with full-body) Data set # Images Faces Upper-body Full-body Availability Face2BMI [5] 4206 ✓ - - ✓ FIW-BMI [33] 7930 ✓ - - - Morph-I [30] 29033 ✓ - - - Dantcheva et al. [4] 1026 ✓ - - ✓ Jiang et al. [6] 5900 - - ✓ - ITU-BMI_F 3372 ✓ - - ✓ ITU-BMI_U 5300 - ✓ - ✓ ITU-BMI_B 537 - - ✓ ✓ ITU-BMI 6105 ✓ ✓ ✓ ✓ I-E Visual-body-to-BMI Jian et al. [6] collected this dataset from ‘reddit.com’ consisting of frontal full-body images. The dataset contains 5900 images of 2950 subjects along with gender, height, and weight labels. Each subject contains a pair of images showing past and present height, and weight. In terms of BMI, the under-weight category has 46 images, 1416 are normal, 1863 are overweight, and 2575 are in obese category. This data set lacks diversity of poses and partial body images because it contains only frontal full-body images. This dataset is not publicly available. IV Proposed Body2BMI-ITU Dataset We have collected 6105 daily-life images from two famous websites ‘w.reddit.com’ and ‘imgur.com’. At these websites, people share their images to show weight loss or weight gain progress to inspire others. In our dataset, each image is annotated with gender, weight, height and BMI information. As show in Figure 2, our dataset contains diverse attributes consisting of selfies, camera pictures, frontal body images, back body images, side poses, only faces, full-body images, half body images (include hip joints). The diversities in poses include people sitting, lying, wearing heavy clothes and caps, may have occluded their faces through cartoons or mobile phones, blackened whole faces or eyes only, may carry food items, pets, and other unavoidable objects in the image. a (a) . b (b) . c (c) . d (d) . e (e) . f (f) . g (g) . h (h) . i (i) . j (j) . Figure 2: Sample images from ITU-BMI dataset: (a) examples with leg’s occlusion. (b) missing faces, (c) sitting position makes height estimation challenging, (d) side-poses, (e) full body images with wide pose variations, (f) holding objects, (g) some outfits make weight estimation more challenging, (h) frontal and side-pose upper-body images (correspond to ITU-BMI_U data set), (i & j) face images only (correspond to ITU-BMI_F data set) To further illustrate the diversity in our dataset, we have separately shown height, weight and BMI distribution of our collected dataset in Figure 3. In figure 2(a) height of subjects ranges between 1.40 meters to 2.20 meters, showing peak values from 1.63 m to 1.73 m. In figure 2(b), four classes of weights are shown ranging from 34 kgs to 250kgs. In figure 2(c) colored bars show four BMI classes including underweight (BMI ≤ 18.5), normal (18.5 << BMI ≤ 25), overweight (25 << BMI ≤ 30), and obese (BMI >> 30). (a) (b) (c) Figure 3: Statistics of the proposed ITU-BMI dataset: (a) distribution of person heights in meters, (b) distribution of person weights in kilograms, (c) distribution of BMI. The dataset collection becomes challenging if the user uploads one image containing both weight loss or weight gain pictures as shown in Fig. 4. In such cases, the target person had to be found manually employing human detection. One such instance is shown Fig. 4a. & 4b. in which uploaded picture contained two or more images of the same person. We use Masked RCNN [37] to separate the humans and then manually corrected the annotation. As shown in Fig. 4b, additional persons also appear in the uploaded pictures and common person is found manually. Other challenges include manual correction of data format and correcting the incorrect correspondence between image and labels. Finally, since different people prefer different units of measurement e.g., kilograms or pounds for weight and centimeters or feet for height, we make all ground consistent: kilograms for weight and meters for height.Because the collected data is posted at a health forum, the uploaded images are with exact measurements to keep track of a person’s success. A person has to register himself to make an account at that forum, and then post their success journey. This forum also includes other suggestions, like keto diet , and everyone is open to share his story, receive comments and chat with others. People measure and upload the exact BMI at the start of their journey to a healthy body with a collage of another taken at some progress point. IV-A Data for Deep Learning We collected a variety of RGB images, to get deep features on them and train a model. However, the data was split into different categories such as only faces, half-body images, and full-body images to apply suitable deep learning models for each category. since, previously there are no trained networks for human body images captured in the wild, we used ResNet-50, and DenseNet for the estimation. For face images, we have used the tained weights of VGG-face to initialize the model and then learned on the wild face images. (a) (b) Figure 4: ITU-BMI dataset cleaning: Mask-RCNN [37] is used to crop the persons however, the cropped persons may still hold other objects with varying backgrounds. (a): a collage of two images of the same person is provided with ‘before’ and ‘after’ weights. (b): a collage of two persons is shown while corresponding data is for only one person. The required common person is manually selected. Figure 5: Image resizing should not alter the ratio of height to width of a person in the image. Therefore, the larger image dimension is re-scaled to the required size while preserving aspect-ratio. The smaller image dimension is then padded with zeros to make it compatible with input layer of deep neural networks. V Proposed Approach We propose the use of deep neural network for the purpose of human body height, weight and BMI estimation using 2D images captured in the wild. We explore different modalities including RGB, gray scale, pose-affinity, depth maps and foreground/background masks. We also explore different architectures including single task and multi-task learning. Figure 6: Proposed deep neural network architectures for estimation of person weight, height and BMI using single-task learning: (a) Single Task Regression (STR), (b) Single Task Classification (STC). Figure 7: Proposed deep neural network architectures for estimation of person weight, height and BMI using multi-task learning: (a) Multi-Task Regression (MTR), (b) Multi-Task Classification (MTC). Figure 8: Sample RGB images from ITU-BMI dataset and corresponding affinity maps computed by Openpose [17]. Pose estimation errors can be seen in some cases. V-A Single-task learning We performed an end-to-end training of ResNet-50 with pre-trained image-net weights and channel-wise normalization. ResNet-50 uses residual links and skip connections to deal with vanishing and exploding gradients problems. Our data set includes large diversity, therefore we choose to fine tune all layer’s weights using low learning rate. The target variables are normalized with their corresponding mean and variance because the units and ranges of height and weight are quite different. As shown in Figure 5a, in the single task learning with regression network (STR), we used mean squared error loss between ground-truth values vitv_i^t and predicted estimations v^it v_i^t both normalized to unit variance and zero mean: vit=(v¯it−μt)/σtv_i^t=( v_i^t- _t)/ _t, where v¯it v_i^t is the observed value, μt _t and σt _t are the mean and variance computed over the training data, and t∈height,weight,BMIt∈\height,weight,BMI\. rt=1m∑i=1m‖vit−v^it‖22,r_t= 1m _i=1^m||v_i^t- v_i^t||_2^2, (1) where m are the number of images in a batch. For single task classification network (STC), as shown in Figure 5.b, we used softmax activation with categorical cross entropy loss function: ft=−1m∑i=1m∑c=1Cℓi,ctlog(ℓ^i,ct),f_t=- 1m _i=1^m _c=1^C _i,c^t ( _i,c^t), (2) where ℓi,c _i,c and ℓ^i,c _i,c are the ground truth and predicted labels for the i-th instance of c-th class, t∈height,weight,BMIt∈\height,weight,BMI\, and C are the total number of classes in our dataset. V-B Multi-task learning The weight, height and BMI of a person are highly correlated to each other, which gives the intuition of learning these parameters simultaneously. A multi-task learning arrangement will facilitate common information sharing across these parameters resulting in improved network performance. Multitask learning shares representations between related tasks by optimizing more than one loss functions at the same time. Since different tasks have different noise patterns, a multitask learning model effectively optimizes all tasks simultaneously and learns a more general representation. Multitask learning improves generalization as it applies the knowledge acquired by learning other correlated tasks. In the current work, we propose a hard parameter sharing configuration of multitask learning, which reduces the risk of over-fitting. In such configuration, we share the initial hidden layers between all tasks, while keep several later layers as task-specific. The shared layers learn the common information across all tasks while the specific layers learn the individual behaviour of each task. We have implemented different multitask learning configurations including multiple regressions and classifications as shown in Figure 6. In case of multi-task regression (MTR), as shown in Figure 6.a, different regression losses are combined: ℒr=λ1rh+λ2rw+λ3rbmi,L_r= _1r_h+ _2r_w+ _3r_bmi, (3) where rhr_h, rwr_w, and rbmir_bmi are the regression losses of height, weight and BMI respectively as defined by Eq. (1), and λ1 _1, λ2, _2, and λ3, _3, are hyper-parameters. For multi-task classification (MTC) as shown in Figure 6.b, the classification losses are combined together. ℒc=λ4fh+λ5fw+λ6fbmi,L_c= _4f_h+ _5f_w+ _6f_bmi, (4) where fhf_h, fwf_w, and fbmif_bmi are the classification losses of height, weight and BMI respectively as defined by Eq. (2), and λ4 _4, λ5 _5 and λ6, _6, are hyper-parameters. V-C Learning Using Different Modalities In addition to the RGB images, we extend our approach to include human body and depth information to facilitate the network in estimating body weight, height and BMI. Body pose estimates the location of body joints which help to better capture the relationship between different body parts of the same person. V-C1 Human Body Parts Affinity Maps Part affinity maps learn association between different human body parts in a given image. To compute these maps we used ’open pose’ proposed by Cao et al. [17]. Open-pose is a multi-stage CNN that produces two different outputs, including confidence maps of different body part locations such as left and right ears, wrists, knees, ankles, etc. The second output is affinity maps showing a degree of association between these body parts. Openpose detects 18 different body joints and by joining nearby joints 19 limbs are obtained. In the current work, we employed convolutional neural networks trained on COCO dataset [60] to obtain parts affinity fields in the proposed dataset, as shown in Fig. 8. The obtained affinity maps are not always accurate, especially on the side pose images and persons holding objects as well as partial occlusions. Therefore, affinity maps alone are not enough for accurate weight, height and BMI estimation. V-C2 Human Body Depth Maps: Depth maps provide information about the body shape of a person. Since there is no human body depth estimation network present to date that is trained on in the wild human body images, so we used NYU-v2 depth pre-trained network [38] for our experimentation. To compute the depth maps we employ encoder-decoder network by AlHashim et al., [38]. In their architecture, DenseNet-169 is used as an encoder, and a series of up-sampling layers are used as a decoder network. The input to the network is RGB image and output is the final depth map at half the input image resolution. The loss function for depth estimation is the summation of ℓ1 _1 error between ground-truth and the network prediction, and also between gradients of the computed and predicted depth maps. In addition, the loss of structural similarity between these depth maps is also included in the objective function. ℒd=fd+β1fg+β2fs,L_d=f_d+ _1f_g+ _2f_s, (5) where β1 _1 and β2 _2 are hyper parameters, fdf_d, fgf_g, and fsf_s are different types of losses defined as: fd=1n∑i=1n|di−d^i|,f_d= 1n _i=1^n|d_i- d_i|, (6) fg=1n∑i=1n|gx(di)−gx(d^i)|+|gy(di)−gy(d^i)|,f_g= 1n _i=1^n|g_x(d_i)-g_x( d_i)|+|g_y(d_i)-g_y( d_i)|, (7) fs=1−SSIM(d,d^)2,f_s= 1-SSIM(d, d)2, (8) where gx(⋅)g_x(·), gy(⋅)g_y(·) are the depth image gradients along X-axis and Y-axis, SSIM(⋅)SSIM(·) is the structural similarity [59] used for image reconstruction tasks. The network is trained on the NYU v2 dataset comprising of various indoor scenes’ videos recorded by both RGB and Depth cameras at a resolution of 640×480640× 480. The dataset contains 120K training samples, and the network is trained on a 50K subset. This dataset also contains some human body images. Obtained depth maps on our proposed ITU-BMI dataset are shown in Fig. 9. Because NYU dataset contains fewer human images, the obtained depth map results are not very accurate. Fig. 9 shows that body parts closer to the camera appear darker and the distant body parts appear lighter in colour. Figure 9: Sample RGB images and the corresponding depth maps computed from NYU depth pre-trained network [38]. Despite some errors, body parts closer to the camera appear dark and body parts distant from camera appear light showing depth variation across the human body. V-C3 Human Body Foreground Masks To damper the effect of background noise and make our network focused on the human body, we have employed human foreground-background segmentation. For this purpose, we used Mask-RCNN [37] which is an instance segmentation deep neural network that employs a region proposal network (RPN), which uses anchors to generate object proposals. Anchors are sets of bounding boxes with predefined locations and scales relative to an input image. All ground-truth classes and bounding boxes are assigned to individual anchors according to a criterion based on Intersection over Union (IoU). In the next stage, another neural network takes proposals from RPN and locates relevant areas on feature map using RoIAlign. This stage outputs object classes, bounding boxes, and a network branch also generates masks for each object on pixel level. Few output masks for our proposed dataset by MaskRCNN are shown in Fig. 10. Our dataset is quite diverse and it contains low-resolution images, therefore the produced masks are not very accurate in some cases. Figure 10: Sample RGB images with corresponding segmented person maps using He et al. [37] on ITU-BMI dataset. In some cases, mask estimation errors can be seen. V-C4 Combining different Modalities: GAD and GAM We have performed experiments with two different combinations of the above discussed modalities. In the first combination, we consider fusion of Gray-scale images with Affinity maps and Depth maps which we dub as ‘GAD’. In the second combination, we fuse Gray-scale, Affinity maps and body Masks dubbed as ‘GAM’. In GAD and GAM, different modalities are concatenated as three channels and the resulting data-structure is input to the proposed network. V-D Transformers for BMI Classification and Regression V-E Transformers for MTC and MTR of BMI, Weight and Height VI Experiments and Results VI-A Experimental Setup All the images in our proposed ITU-BMI dataset are resized to 224×224224× 224 while maintaining the aspect ratio to ensure compatibility with the input layers of the feature extractor networks which are VGG16, ResNet50, and Denset121. For the purpose of resize, the larger image dimension is first re-scaled to 224 and the smaller dimension is then zero-padded to obtain the required size (Fig. 5). Our dataset contains 6105 images which are randomly divided into 80/20 train/test split. The ITU-BMI dataset contains 537 full body images (taken from test set only) referred to as ITU-BMI_B dataset and 5300 images contain only upper body referred as ITU-BMI_U dataset. Openpose face detector is used to crop faces in the dataset resulting in 3372 facial images referred as ITU-BMI_F dataset. More details of these categories are shown in Table I. We used ResNet-50 as baseline architecture with our proposed STR, STC, MTR, MTC networks (Fig. 5, 6) to predict height, weight, and BMI of a person from his given image. In each network, we added 2 dense fully connected layers of 128 and 32 neurons with leakyrelu activation function (alpha=0.3) to the ResNet50 baseline model. After that three separate dense layers are used in case of multitask regression/ classification, and one dense layer is used for single-task regression/classification. For initialization, Imagenet pre-trained weights were used, followed by complete end-to-end training. We used stochastic gradient descent (SGD) optimizer with learning rate = 0.001, decay rate= 0.0008, batch size =16, and number of epochs ≥ 100. A weight and bias decay of 5e−45e^-4 was also applied to all layers of the model. In the current work, the hyper-parameters used in the equation 3, and 4 are empirically found. More details will be discussed in Ablation study section. VI-B Regression Experiments We provide a complete analysis of different learning techniques to predict height, weight, and BMI of a human from 2D body image. Mean absolute error (MAE) and accuracy (%Acc.) metrices are being used to evaluate regression and classification results. We have compared the proposed approach with existing state of the art (SOTA) methods on the proposed dataset. Comparison of Single-task Vs. Multi-task Regression: Table I shows the comparison of MTR with STR and current state of the art including Dantcheva et al. [4] and Jiang et al. [6]. The proposed MTR has consistently shown lower MAE over all compared methods on all four proposed datasets. The STR has produced the second best results which are also significantly better than the existing state of the art methods. For the case of MTR, the BMI regression over face dataset (ITU-BMI_F) has obtained an error of 4.87 which is significantly higher than the BMI estimated from the full body (ITU-BMI_B) which is 3.32. ITU-BMI dataset which includes upper body, and fully body has obtained 3.73 MAE which is lesser than the upper body BMI due to more training images and full body images. For the weight estimation, minimum MAE of 12.60 Kg is obtained by MTR on ITU-BMI_B which is significantly lower than the weight estimation from face images in ITU-BMI_f, which is 17.11 Kg. For height estimation, minimum MAE of 0.08m is observed on ITU-BMI_B dataset which is better compared to ITU-BMI_F and ITU-BMI_U images. We observe that full human body images provide more information, and produce more accurate estimations of H, W and BMI as compared to the upper-body or face images only. TABLE I: Performance comparison of the proposed Single-task Regression (STR), Multi-task Regression (MTR) for BMI, weight (kilograms), and height (meters) estimation, in terms of Means Absolute Error (MAE) on the proposed data sets using RGB images Datasets Task STR MTR Dantcheva et al. Jiang et al. [6] [4] SVR GPR ITU-BMI_F BMI 5.21 4.87 7.13 - - W 17.62 17.11 22.59 - - H 0.11 0.09 0.44 - - ITU-BMI_U BMI 4.17 4.11 5.94 - - W 15.11 14.89 19.58 - - H 0.11 0.09 0.38 - - ITU-BMI_B BMI 4.39 3.32 5.33 12.49 12.66 W 14.92 12.60 17.63 23.77 24.04 H 0.09 0.08 0.36 0.10 0.09 ITU-BMI BMI 4.29 3.73 5.86 - - W 14.82 13.69 19.11 - - H 0.11 0.08 0.37 - - Among the existing BMI prediction methods, Jiang et al. [6], similar to our approach estimates BMI using only fully body images. However, the code and the dataset used by them is not publicaly available. In order to report their performance on ITU-BMI_B dataset, we replicated their method by closely following the approach mentioned in their paper. The results of Jiang et al. using SVR and GPR were compared with our approaches in Table I. Both versions of Jiang et al. has produced significantly larger error for the case of BMI and weight prediction. For the case of height prediction where GPR version has obtained the same MAE as the proposed STR approach. Their degraded performance may be attributed to the lack of representation power of anthropometric features which require frontal full body images. In our dataset, the pose varies widely therefore, the anthropometric features have performed poor. The comparison has also been made with the work of Dantcheva et al. [4] which was originally proposed for BMI, weight and height prediction using only frontal face images. In our proposed ITU-BMI_F dataset, facial pose varies significantly which has resulted in the degraded performance of their algorithm. In addition to the face images, their algorithm is also trained and evaluated for the uppper body, fully body and the overall ITU-BMI dataset. On full body, Dantcheva et al. has performed significantly better than Jiang et al. and remian a close competitor to the STR network. However, the performance of the proposed MTR network has remained the best across all the experiments. Figure 11: Sample RGB images with overlaid body joint positions as detected by Openpose [17]. Missing joints can be observed in some cases resulting in height, weight and BMI estimation errors. Comparison of GAD and GAM Modalities. In addition to the RGB images, the proposed STR and MTR methods have also been evaluated on Gray scale-Affinity-Depth (GAD) and Gray scale-Affinity-Mask (GAM) multi-modality images as discussed in Section V.C. The experiments are performed on full ITU-BMI dataset using proposed STR and MTR networks and results are shown in Table I. We observe that in all the experiments GAM has resulted better performance than GAD which is mainly due to the reduce quality of depth images. A better estimation method may improve the GAD results. We also observe that MTR is able to obtain performance than STR which may be attributed to the overlap information between parameters to be estimated. The proposed approach is also compared with Dantcheva et al. [4] method. However, the performance of MTR remains the best in all experiments. These findings are consistent with the observations in the Table I. TABLE I: MAE results on proposed ITU-BMI data set using GAD& GAM images Modalities Task STR MTR Method [4] GAD BMI 4.70 4.39 6.24 W 17.21 15.91 20.53 H 0.11 0.10 0.48 GAM BMI 4.01 4.11 6.46 W 15.56 15.11 21.46 H 0.11 0.09 0.51 Analysis of MTR across different BMI classes: The MTR architecture is found to be the best performer in a wide range of experiments. Therefore, we further analyze only MTR results across different BMI classes as defined in Section IV. Table IV demonstrates that full body images dataset ITU-BMI_B has resulted in the minimum MAE across all the classes. The performance comparison of different BMI classes show maximum MAE in obese and underweight classes which is due to the minimum number of training examples in these classes. For the case of normal and overweight classes, the error is reduced due to more training examples in these classes. TABLE IV: A comparison of MAE amongst different BMI classes using MTR. Datasets All Underweight Normal Overweight Obese ITU-BMI 3.73 3.82 2.95 2.76 4.93 ITU-BMI_U 5.94 4.38 3.09 2.80 5.47 ITU-BMI_F 4.87 7.97 3.99 3.09 5.97 ITU-BMI_B 3.32 3.34 2.43 2.36 4.74 Analysis of MTR across different backbone networks: In addition to the ResNet-50 network, which is used in the all previous experiments, we have also done some experimentation using VGG-16 [21] and DenseNet-121 [58] as backbone architecture in the proposed MTR approach. The experimental results in the Table V show that ResNet-50 has resulted in the better estimation of BMI, H, and W as compared to the VGG-16 and DenseNet-121. TABLE V: Multitask Regression MAE results using different baseline architectures on ITU-BMI RGB images Network BMI Weight[kg] Height[m] VGG16 4.23 14.73 0.08 DenseNet121 4.11 14.56 0.09 ResNet50 3.73 13.69 0.08 VI-C Classification Experiments In the previous subsection, the values of BMI, weight and height were estimated using deep regression networks. Based on different BMI values, underweight, normal, overweight and obese categories are defined in the literature as shown in the Table VI. We have also defined four classes for height and weight as well. Comparison of Single-task Vs. Multi-task Classification for Different Modalities: The images in the proposed ITU-BMI dataset are grouped into different classes based on height and weight ranges. For each task, images are divided into four different classes, and the ranges are selected such that approximately balanced number of samples are obtained in each class as shown in Table VI. For the case of BMI, the four classes are defined as per World Health Organisation (WHO) categorization. We performed single-task and multi-task classification on RGB, GAD, and GAM images and compare the classification accuracy (% Acc.), and Area Under the Curve (%AUC.) for different modalities. TABLE VI: Summary of Height (meters), Weight (kilograms) and BMI classes Class 0 Class 1 Class 2 Class 3 Height [m] 1.30-1.64 1.65-1.71 1.72-1.81 ≥ 1.82 # Samples 1718 1567 1555 1265 Weight [kg] 34.0-64.3 64.4-83.7 83.8-103.1 ≥ 103.2 # Samples 1508 1624 1408 1565 BMI 10-18.5 18.6-25.0 25.1- 30.0 ≥ 30.1 # Samples 731 1478 1472 2424 For BMI classification, we define four classes, where class 0 is underweight, class 1 is normal, class 2 is overweight and class 3 is obese. All classification results are reported on a randomly selected hold-out 20% test and 80% train dataset. Figure. 5b. and 6b. show the STC and MTC network architectures, respectively. Table VII shows the classification results and Fig. 12, 13, and 14 show the corresponding confusion matrices of height, weight, and BMI based classification using STC and MTC employing different modalities. From Table VII, we can see that multi-task classification (MTC) has performed better than single-task classification (STC) in all experiments. In case of BMI, the best performance is obtained on GAM modaility which is 64.61% classification accuracy and 80.9% AUC. by MTC. AUC measure tells sensitivity which means how well predictions are ranked, rather than their absolute values. AUC provides an aggregate measure of performance across all possible classification thresholds. Classification-threshold invariance is useful in this optimization problem, because it is not critical to minimize one type of classification error. It means there does not exist much disparity between the cost of false negatives vs. false positives. The heighest correlation is found between height and BMI which is 0.80, and between weight and BMI is 0.58, which is less and does not suggest a linear machine learning model to be fitted. This observation led us to use deep learning feature extraction to be used for estimations. Using deep features multitask predictions of height, weight, and BMI as dependent variables provide more accuracy compared to single-task predictions. For the case of Weight and Height, RGB modalities has produced the best results. However, the GAD has remained the second best for the case of weight based classification and GAM has remained second best for the case of height based classification. TABLE VII: Classification accuracy (%Acc.) and Area under Curve (%AUC) of H, W & BMI on proposed ITU-BMI data set using ResNet50 Modalities BMI W H Acc. AUC Acc. AUC Acc. AUC RGB STC 62.32 80. 6 57.82 80.9 47.82 71.5 MTC 62.08 80.3 58.39 81.1 50.36 73.8 GAD STC 60.27 78.4 55.03 77.9 46.68 72.3 MTC 64.12 80.7 58.06 79.9 46.19 73.3 GAM STC 61.26 77.5 54.95 78.5 47.50 73.1 MTC 64.61 80.9 57.00 78.0 47.82 73.2 (a) (b) (c) (d) (e) (f) Figure 12: Confusion matrices of STC and MTC on RGB images. Subfigures (a), (b) and (c) correspond to BMI, W and H based classification using STC. Subfigures (d), (e) and (f) correspond to BMI, W and H based classification using MTC. (a) (b) (c) (d) (e) (f) Figure 13: Confusion matrices of STC and MTC on GAD images. Subfigures (a), (b) and (c) correspond to BMI, W and H based classification using STC. Subfigures (d), (e) and (f) correspond to BMI, W and H based classification using MTC. (a) (b) (c) (d) (e) (f) Figure 14: Confusion matrices of STC and MTC on GAM images. Subfigures (a), (b) and (c) correspond to BMI, W and H based classification using STC. Subfigures (d), (e) and (f) correspond to BMI, W and H based classification using MTC. Analysis of MTC across different backbone networks: We have also experimented MTC on RGB images using different deep-nets as baseline architectures to compare the performance as shown in Table VIII. The DenseNet-121 has relatively performed better than other two backbone networks, however, the performance difference is not significant. Motivated by our experiments on MTR, where ResNet-50 was the best performer, most of the experiments reported in this work use ResNet-50 as a backbone. TABLE VIII: Multitask classification Acc. and AUC (%) of H, W, & BMI on RGB images using different deep nets. Task Networks Acc. AUC BMI ResNet-50 62.40 80.5 DenseNet-121 62.08 80.8 VGG-16 62.81 80.4 W ResNet-50 57.16 80.0 DenseNet-121 59.54 79.7 VGG-16 59.54 77.7 H ResNet-50 48.48 72.2 DenseNet-121 50.20 72.6 VGG-16 48.07 68.7 VII Discussion We have used mean absolute error (MAE) to estimate regression error. MAE works for same scale of data, and measures the absolute difference between two continuous variables. Our MAE values show the average difference between predicted and target values of H, W, and BMI of ITU-BMI dataset. In general, mean squared error (MSE) is also used to measure the effectiveness of a regression model. VIII Conclusion In this work, a new dataset and deep learning based framework is proposed for robust estimation of weight, height, and Body Mass Index (BMI) using images captured in uncontrolled environments. For this purpose, deep neural networks are used for regression over weight, height and BMI. In addition to that, based on weight, height and BMI, the dataset is labelled into five categories including under-weight, normal, over-weight, and obese using the WHO recommendations. Deep neural networks are also trained for the purpose of classification using visual images as input. For both regression and classification, single task as well as multitask approaches are compared and multitask approaches are found to be the more accurate. In addition to RGB mode, other modalities such as depth, edge masks, and affinity maps are also proposed to get improved performance. Extensive experiments are performed using various backbone CNNs including ResNet-50, DesNet-121, and VGG-16. In most of the experiments, ResNet-50 has yielded the best performance. A new dataset consisting of 6105 images with weight and height labels, extracted features, and the trained networks will soon be made publicly available for research purposes. References [1] Dantcheva, Antitza and Elia, Petros and Ross, Arun. What else does your biometric data reveal? A survey on soft biometrics. IEEE Transactions on Information Forensics and Security, volume 11, number 3, pages 441–467. IEEE, 2015. [2] Gonzalez-Sosa, Ester and Fierrez, Julian and Vera-Rodriguez, Ruben and Alonso-Fernandez, Fernando. Facial soft biometrics for recognition in the wild: Recent works, annotation, and COTS evaluation. IEEE Transactions on Information Forensics and Security, volume 13, number 8, pages 2001–2014. IEEE, 2018. [3] Neal, Tempestt J and Woodard, Damon L. You are not acting like yourself: A study on soft biometric classification, person identification, and mobile device use. IEEE Transactions on Biometrics, Behavior, and Identity Science, volume 1, number 2, pages 109–122. IEEE, 2019. [4] Dantcheva, Antitza and Bremond, Francois and Bilinski, Piotr. Show me your face and I will tell you your height, weight and body mass index. 2018 24th International Conference on Pattern Recognition (ICPR), pages 3555–3560. IEEE, 2018. [5] Kocabey, Enes and Camurcu, Mustafa and Ofli, Ferda and Aytar, Yusuf and Marin, Javier and Torralba, Antonio and Weber, Ingmar. Face-to-bmi: Using computer vision to infer body mass index on social media. Eleventh International AAAI Conference on Web and Social Media, 2017. [6] Jiang, Min and Guo, Guodong. Body Weight Analysis from Human Body Images. IEEE Transactions on Information Forensics and Security. IEEE, 2019. [7] Bell, Dane and Laparra, Egoitz and Kousik, Aditya and Ishihara, Terron and Surdeanu, Mihai and Kobourov, Stephen. Detecting Diabetes Risk from Social Media Activity. Proceedings of the Ninth International Workshop on Health Text Mining and Information Analysis, pages 1–11, 2018. [8] Arnold, Melina and Leitzmann, Michael and Freisling, Heinz and Bray, Freddie and Romieu, Isabelle and Renehan, Andrew and Soerjomataram, Isabelle. Obesity and cancer: an update of the global impact. Cancer epidemiology, volume 41, pages 8–15. Elsevier, 2016. [9] Lewis, Matthew T and Lujan, Heidi L and Tonson, Anne and Wiseman, Robert W and DiCarlo, Stephen E. Obesity and inactivity, not hyperglycemia, cause exercise intolerance in individuals with type 2 diabetes: Solving the obesity and inactivity versus hyperglycemia causality dilemma. Medical hypotheses, volume 123, pages 110–114. Elsevier, 2019. [10] Ling, Charlotte and Rön, Tina. Epigenetics in human obesity and type 2 diabetes. Cell metabolism. Elsevier, 2019. [11] Narayanan, Ajay and Yogesh, Ahana and Mitchell, Ron B and Johnson, Romaine F. Asthma and obesity as predictors of severe obstructive sleep apnea in an adolescent pediatric population. The Laryngoscope. Wiley Online Library, 2019. [12] Xu, Shujing and Gilliland, Frank D and Conti, David V. Elucidation of causal direction between asthma and obesity: a bi-directional Mendelian randomization study. International journal of epidemiology, 2019. [13] Uppunda, Deepak and Shetty, Ranjan K and Rao, Pragna and Razak, Abdul and Shetty, Kiran and Chauhan, Sheetal and Singh, Ajit and others. Association of Metabolic Obesity and BMI Status with Severity of Angiographic Coronary Artery Disease in Stable Angina Patients.. Journal of Clinical & Diagnostic Research, volume 13, number 4, 2019. [14] Kanazawa, Ippei and Notsu, Masakazu and Takeno, Ayumu and Tanaka, Ken-ichiro and Sugimoto, Toshitsugu. Overweight and underweight are risk factors for vertebral fractures in patients with type 2 diabetes mellitus. Journal of bone and mineral metabolism, volume 37, number 4, pages 703–710. Springer, 2019. [15] Ruffner, Melanie A and Sullivan, Kathleen E and others. Complications associated with underweight primary immunodeficiency patients: prevalence and associations within the USIDNET Registry. Journal of clinical immunology, volume 38, number 3, pages 283–293. Springer, 2018. [16] Nakagawa, Yuichi and Nakanishi, Toshiki and Satake, Eiichiro and Matsushita, Rie and Saegusa, Hirokazu and Kubota, Akira and Natsume, Hiromune and Shibata, Yukinobu and Fujisawa, Yasuko. Postnatal BMI changes in children with different birthweights: A trial study for detecting early predictive factors for pediatric obesity. Clinical Pediatric Endocrinology, volume 27, number 1, pages 19–29. The Japanese Society for Pediatric Endocrinology, 2018. [17] Cao, Zhe and Hidalgo, Gines and Simon, Tomas and Wei, Shih-En and Sheikh, Yaser. OpenPose: realtime multi-person 2D pose estimation using Part Affinity Fields. arXiv preprint arXiv:1812.08008, 2018. [18] Wen, Lingyun and Guo, Guodong. A computational approach to body mass index prediction from face images. Image and Vision Computing, volume 31, number 5, pages 392–400. Elsevier, 2013. [19] Velardo, Carmelo and Dugelay, Jean-Luc. Weight estimation from visual body appearance. 2010 Fourth IEEE International Conference on Biometrics: Theory, Applications and Systems (BTAS), pages 1–6. IEEE, 2010. [20] Parkhi, Omkar M and Vedaldi, Andrea and Zisserman, Andrew and others. Deep face recognition.. bmvc, volume 1, number 3, pages 6, 2015. [21] Simonyan, Karen and Zisserman, Andrew. Very deep convolutional networks for large-scale image recognition. arXiv preprint arXiv:1409.1556, 2014. [22] Ozbulak, Gokhan and Aytar, Yusuf and Ekenel, Hazim Kemal. How transferable are CNN-based features for age and gender classification?. 2016 International Conference of the Biometrics Special Interest Group (BIOSIG), pages 1–6. IEEE, 2016. [23] Nahavandi, D and Abobakr, A and Haggag, H and Hossny, M and Nahavandi, S and Filippidis, D. A skeleton-free kinect system for body mass index assessment using deep neural networks. 2017 IEEE International Systems Engineering Symposium (ISSE), pages 1–6. IEEE, 2017. [24] National Health and Nutrition Examination Survey, Center for Disease. Control and Prevention, 2006. [25] Pfitzner, Christian and May, Stefan and Nüchter, Andreas. Evaluation of Features from RGB-D Data for Human Body Weight Estimation. IFAC-PapersOnLine, volume 50, number 1, pages 10148–10153. Elsevier, 2017. [26] Kocabey, Enes and Ofli, Ferda and Marin, Javier and Torralba, Antonio and Weber, Ingmar. Using computer vision to study the effects of bmi on online popularity and weight-based homophily. International Conference on Social Informatics, pages 129–138. Springer, 2018. [27] Barr, Makenzie and Guo, Guodong and Colby, Sarah and Olfert, Melissa. Detecting body mass index from a facial photograph in lifestyle intervention. Technologies, volume 6, number 3, pages 83. Multidisciplinary Digital Publishing Institute, 2018. [28] Drucker, Harris and Burges, Christopher JC and Kaufman, Linda and Smola, Alex J and Vapnik, Vladimir. Support vector regression machines. Advances in neural information processing systems, pages 155–161, 1997. [29] Viola, Paul and Jones, Michael. Robust real-time face detection. null, pages 747. IEEE, 2001. [30] Ricanek, Karl and Tesafaye, Tamirat. Morph: A longitudinal image database of normal adult age-progression. 7th International Conference on Automatic Face and Gesture Recognition (FGR06), pages 341–345. IEEE, 2006. [31] Milborrow, Stephen. Locating facial features with active shape models. University of Cape Town, 2007. [32] Jahandideh, Rashidedin and Targhi, Alireza Tavakoli and Tahmasbi, Maryam. Physical Attribute Prediction Using Deep Residual Neural Networks. arXiv preprint arXiv:1812.07857, 2018. [33] Jiang, Min and Shang, Yuanyuan and Guo, Guodong. On visual BMI analysis from facial images. Image and Vision Computing, volume 89, pages 183–196. Elsevier, 2019. [34] Wu, Xiang and He, Ran and Sun, Zhenan and Tan, Tieniu. A light cnn for deep face representation with noisy labels. IEEE Transactions on Information Forensics and Security, volume 13, number 11, pages 2884–2896. IEEE, 2018. [35] Wen, Yandong and Zhang, Kaipeng and Li, Zhifeng and Qiao, Yu. A discriminative feature learning approach for deep face recognition. European conference on computer vision, pages 499–515. Springer, 2016. [36] Deng, Jiankang and Guo, Jia and Xue, Niannan and Zafeiriou, Stefanos. Arcface: Additive angular margin loss for deep face recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pages 4690–4699, 2019. [37] He, Kaiming and Gkioxari, Georgia and Dollár, Piotr and Girshick, Ross. Mask r-cnn. Proceedings of the IEEE international conference on computer vision, pages 2961–2969, 2017. [38] Alhashim, Ibraheem and Wonka, Peter. High quality monocular depth estimation via transfer learning. arXiv preprint arXiv:1812.11941, 2018. [39] Henderson, Audrey J and Holzleitner, Iris J and Talamas, Sean N and Perrett, David I. Perception of health from facial cues. Philosophical Transactions of the Royal Society B: Biological Sciences, volume 371, number 1693, pages 20150380. The Royal Society, 2016. [40] Mayer, Christine and Windhager, Sonja and Schaefer, Katrin and Mitteroecker, Philipp. BMI and WHR are reflected in female facial shape and texture: a geometric morphometric image analysis. PloS one, volume 12, number 1, pages e0169336. Public Library of Science San Francisco, CA USA, 2017. [41] Lee, Seon Yeong and Gallagher, Dympna. Assessment methods in human body composition. Current opinion in clinical nutrition and metabolic care, volume 11, number 5, pages 566. NIH Public Access, 2008. [42] Zhu, Yu and Li, Yan and Mu, Guowang and Shan, Shiguang and Guo, Guodong. Still-to-video face matching using multiple geodesic flows. IEEE Transactions on Information Forensics and Security, volume 11, number 12, pages 2866–2875. IEEE, 2016. [43] Günther, Manuel and Hu, Peiyun and Herrmann, Christian and Chan, Chi-Ho and Jiang, Min and Yang, Shufan and Dhamija, Akshay Raj and Ramanan, Deva and Beyerer, Jürgen and Kittler, Josef and others. Unconstrained face detection and open-set face recognition challenge. 2017 IEEE International Joint Conference on Biometrics (IJCB), pages 697–706. IEEE, 2017. [44] Wang, Qiangchang and Guo, Guodong and Nouyed, Mohammad Iqbal. Learning channel inter-dependencies at multiple scales on dense networks for face recognition. arXiv preprint arXiv:1711.10103, 2017. [45] Pfitzner, Christian and May, Stefan and Nüchter, Andreas. Body weight estimation for dose-finding and health monitoring of lying, standing and walking patients based on RGB-D data. Sensors, volume 18, number 5, pages 1311. Multidisciplinary Digital Publishing Institute, 2018. [46] Affuso, Olivia and Pradhan, Ligaj and Zhang, Chengcui and Gao, Song and Wiener, Howard W and Gower, Barbara and Heymsfield, Steven B and Allison, David B. A method for measuring human body composition using digital images. PloS one, volume 13, number 11, pages e0206430. Public Library of Science San Francisco, CA USA, 2018. [47] Andersson, Virginia Ortiz and Amaral, Livia S and Tonini, Aline R and Araujo, Ricardo M. Gender and body mass index classification using a microsoft kinect sensor. The Twenty-Eighth International Flairs Conference, 2015. [48] Farina, Gian Luca and Spataro, Fabrizio and De Lorenzo, Antonino and Lukaski, Henry. A smartphone application for personal assessments of body composition and phenotyping. Sensors, volume 16, number 12, pages 2163. Multidisciplinary Digital Publishing Institute, 2016. [49] Günel, Semih and Rhodin, Helge and Fua, Pascal. What Face and Body Shapes Can Tell Us About Height. 2019 IEEE/CVF International Conference on Computer Vision Workshop (ICCVW), pages 1819–1827. IEEE, 2019. [50] Pascali, Maria Antonietta and Giorgi, Daniela and Bastiani, Luca and Buzzigoli, E and Henríquez, Pedro and Matuszewski, Bogdan J and Morales, M-A and Colantonio, Sara. Face morphology: Can it tell us something about body weight and fat?. Computers in biology and medicine, volume 76, pages 238–249. Elsevier, 2016. [51] Jiang, Min and Guo, Guodong and Mu, Guowang. Visual BMI estimation from face images using a label distribution based method. Computer Vision and Image Understanding, pages 102985. Elsevier, 2020. [52] Vakli, Pál and Deák-Meszlényi, Regina J and Auer, Tibor and Vidnyánszky, Zoltán. Predicting Body Mass Index From Structural MRI Brain Images Using a Deep Convolutional Neural Network. Frontiers in Neuroinformatics, volume 14, pages 10. Frontiers, 2020. [53] Ilao, Adomar L and Cardino, Adrian Christopher and Fernandez, Clarence and Saulon, Lawrence. BMIMatic: Body mass index derivation from captured images. Eleventh International Conference on Digital Image Processing (ICDIP 2019), volume 11179, pages 111790J. International Society for Optics and Photonics, 2019. [54] Pantanowitz, Adam and Cohen, Emmanuel and Gradidge, Philippe and Crowther, Nigel and Aharonson, Vered and Rosman, Benjamin and Rubin, David M. Estimation of Body Mass Index from Photographs using Deep Convolutional Neural Networks. arXiv preprint arXiv:1908.11694, 2019. [55] Verma, Shalini and Hussain, M Ejaz. Obesity and diabetes: an update. Diabetes & Metabolic Syndrome: Clinical Research & Reviews, volume 11, number 1, pages 73–79. Elsevier, 2017. [56] Alpert, Martin A and Karthikeyan, Kamalesh and Abdullah, Obai and Ghadban, Rugheed. Obesity and cardiac remodeling in adults: mechanisms and clinical implications. Progress in cardiovascular diseases, volume 61, number 2, pages 114–123. Elsevier, 2018. [57] Lin, Hsien-Ho and Wu, Chieh-Yin and Wang, Chih-Hui and Fu, Han and Lönnroth, Knut and Chang, Yi-Cheng and Huang, Yen-Tsung. Association of obesity, diabetes, and risk of tuberculosis: two population-based cohorts. Clinical Infectious Diseases, volume 66, number 5, pages 699–705. Oxford University Press US, 2018. [58] Iandola, Forrest and Moskewicz, Matt and Karayev, Sergey and Girshick, Ross and Darrell, Trevor and Keutzer, Kurt. Densenet: Implementing efficient convnet descriptor pyramids. arXiv 2014. arXiv preprint arXiv:1404.1869. [59] Wang, Zhou and Bovik, Alan C and Sheikh, Hamid R and Simoncelli, Eero P. Image quality assessment: from error visibility to structural similarity. IEEE transactions on image processing, volume 13, number 4, pages 600–612. IEEE, 2004. [60] Lin, Tsung-Yi and Maire, Michael and Belongie, Serge and Hays, James and Perona, Pietro and Ramanan, Deva and Dollár, Piotr and Zitnick, C Lawrence. Microsoft coco: Common objects in context. European conference on computer vision, pages 740–755. Springer, 2014. [61] Khan, Soofia and Xanthakos, Stavra A and Hornung, Lindsey and Arce-Clachar, Catalina and Siegel, Robert and Kalkwarf, Heidi J. Relative accuracy of bioelectrical impedance analysis for assessing body composition in children with severe obesity. Journal of pediatric gastroenterology and nutrition, volume 70, number 6, pages e129–e135. LWW, 2020. [62] Vermeiren, Eline and Ysebaert, Marijke and Van Hoorenbeeck, Kim and Bruyndonckx, Luc and Van Dessel, Kristof and Van Helvoirt, Maria and De Guchtenaere, Ann and De Winter, Benedicte and Verhulst, Stijn and Van Eyck, Annelies. Comparison of bioimpedance spectroscopy and dual energy X-ray absorptiometry for assessing body composition changes in obese children during weight loss. European Journal of Clinical Nutrition, volume 75, number 1, pages 73–84. Nature Publishing Group, 2021. [63] Bjorkman, Mikko P and Jyvakorpi, Satu K and Strandberg, Timo E and Pitkala, Kaisu H and Tilvis, Reijo S. The associations of body mass index, bioimpedance spectroscopy-based calf intracellular resistance, single-frequency bioimpedance analysis and physical performance of older people. Aging clinical and experimental research, volume 32, number 6, pages 1077–1083. Springer, 2020. [64] Shannon, Carley A and Brown, Justin R and Del Pozzi, Andrew T. Comparison of Body Composition Prediction Equations with Air Displacement Plethysmography in Overweight and Obese Caucasian Males. International journal of exercise science, volume 12, number 4, pages 1034. Western Kentucky University, 2019. [65] Pellonperä, Outi and Koivuniemi, Ella and Vahlberg, Tero and Mokkala, Kati and Tertti, Kristiina and Rönnemaa, Tapani and Laitinen, Kirsi. Body composition measurement by air displacement plethysmography in pregnancy: Comparison of predicted versus measured thoracic gas volume. Nutrition, volume 60, pages 227–229. Elsevier, 2019. [66] Ashby-Thompson, Maxine and Ji, Ying and Wang, Jack and Yu, Wen and Thornton, John C and Wolper, Carla and Weil, Richard and Chambers, Earle C and Laferrère, Blandine and Pi-Sunyer, F Xavier and others. High-Resolution Three-Dimensional Photonic Scan-Derived Equations Improve Body Surface Area Prediction in Diverse Populations. Obesity, volume 28, number 4, pages 706–717. Wiley Online Library, 2020. [67] Wells, Jonathan CK. Three-dimensional optical scanning for clinical body shape assessment comes of age. The American journal of clinical nutrition, volume 110, number 6, pages 1272–1274. Oxford University Press, 2019. [68] Pasanta, Duanghathai and Tungjai, Montree and Chancharunee, Sirirat and Sajomsang, Warayuth and Kothan, Suchart. Body mass index and its effects on liver fat content in overweight and obese young adults by proton magnetic resonance spectroscopy technique. World journal of hepatology, volume 10, number 12, pages 924. Baishideng Publishing Group Inc, 2018. [69] Bini, Jason and Bhatt, Shivani and Hillmer, Ansel T and Gallezot, Jean-Dominique and Nabulsi, Nabeel and Pracitto, Richard and Labaree, David and Kapinos, Michael and Ropchan, Jim and Matuskey, David and others. Body Mass Index and Age Effects on Brain 11β-Hydroxysteroid Dehydrogenase Type 1: a Positron Emission Tomography Study. Molecular imaging and biology, volume 22, number 4, pages 1124–1131. Springer, 2020. [70] Vaswani, Ashish and Shazeer, Noam and Parmar, Niki and Uszkoreit, Jakob and Jones, Llion and Gomez, Aidan N and Kaiser, Łukasz and Polosukhin, Illia. Attention is all you need. Advances in neural information processing systems, volume 30, 2017. [71] Dosovitskiy, Alexey and Beyer, Lucas and Kolesnikov, Alexander and Weissenborn, Dirk and Zhai, Xiaohua and Unterthiner, Thomas and Dehghani, Mostafa and Minderer, Matthias and Heigold, Georg and Gelly, Sylvain and others. An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929, 2020. [72] Eline Vermeiren, Marijke Ysebaert, Kim Van Hoorenbeeck, Luc Bruyndonckx, Kristof Van Dessel, Maria Van Helvoirt, Ann De Guchtenaere, Benedicte De Winter, Stijn Verhulst, and Annelies Van Eyck. Comparison of bioimpedance spectroscopy and dual energy X-ray absorptiometry for assessing body composition changes in obese children during weight loss. European Journal of Clinical Nutrition, volume 75, number 1, pages 73–84, 2021.